The process of preparing a DNA or RNA sample so it can be compatible with an NGS sequencing instrument. This usually involves fragmenting the nucleic acid and attaching specialized adapters to the ends of the fragments.
Short, chemically synthesized, single-stranded DNA molecules that are ligated (attached) to the ends of DNA fragments during library preparation. They contain specific sequences necessary for anchoring the fragment to the sequencer's flow cell and binding sequencing primers.
A unique, short DNA sequence (usually 6–10 base pairs) embedded within the adapter. By attaching a distinct index to every fragment of a specific sample, scientists can mix dozens of different samples together into a single sequencing run (multiplexing) and later separate the data using software (demultiplexing).
The physical, fluidics-designed glass slide or consumable chip where the actual sequencing reaction occurs. The surface of the flow cell is coated with millions of specialized oligonucleotides (oligos) that bind to the library adapters, immobilizing the DNA fragments for clonal amplification and sequencing.
An in-instrument amplification process (such as bridge amplification or rolling circle amplification) that duplicates a single immobilized DNA fragment into a localized "cluster" of thousands of identical copies. This multiplies the fluorescent or chemical signal, making it strong enough for the sequencer's sensors to detect.
The actual discrete sequence of nucleobases (A, T, C, G) generated by the sequencer from a single DNA fragment. Depending on the technology, reads can be short (50–300 base pairs for Illumina) or long (10,000+ base pairs for PacBio or Oxford Nanopore).
A sequencing method where a single DNA fragment is sequenced from both ends—first from one direction (Read 1), and then from the opposite direction (Read 2). This provides highly accurate alignment data, particularly useful for identifying structural variants or genomic rearrangements.
The bioinformatic process of comparing raw sequencing reads against a known reference genome to determine exactly where in the genome those sequences originated.
The average number of times a specific nucleotide position in the genome is sequenced by independent reads. For example, "30x coverage" means that, on average, every base across the targeted region was read 30 times. Higher depth increases statistical confidence when identifying rare variants.
The standard text-based file format used to store raw data from an NGS run. For every read, a FASTQ file contains the sequence name, the actual sequence string (A, T, C, G), and a corresponding string of characters representing the Phred quality score (the probability that each base call is correct).
The bioinformatic analysis phase that identifies differences between the aligned sequencing data and the reference genome. These variations can include Single Nucleotide Polymorphisms (SNPs), small insertions/deletions (indels), or large structural copy number variations (CNVs).
The universal text file format used to output the results of variant calling. A VCF file lists the chromosome, exact genomic position, reference base, mutated alternative base, and various metadata metrics assessing the quality and frequency of each detected genetic variant.
Copyright © 2026 Singapore , Thailand
ALL RIGHTS RESERVED