DNA sequencing
Determining the order of nucleotides in DNA.
National Institute of Standards and Technology · Public domain
DNA sequencing is the process of determining the nucleic acid sequence – the order of nucleotides in DNA. It includes any method or technology used to determine the order of the four bases: adenine, thymine, cytosine, and guanine. The advent of rapid DNA sequencing methods has greatly accelerated biological and medical research and discovery.
Lore & Background
The first DNA sequences were obtained in the early 1970s by academic researchers using laborious methods based on two-dimensional chromatography. Following the development of fluorescence-based sequencing methods with a DNA sequencer, DNA sequencing became easier and orders of magnitude faster. Knowledge of DNA sequences has become indispensable for basic biological research, DNA Genographic Projects, and in numerous applied fields such as medical diagnosis, biotechnology, forensic biology, virology, and biological systematics.
Reader's Guide
DNA sequencing has become a key technology in many areas of biology and other sciences. It allows researchers to identify changes in genes and noncoding DNA, associations with diseases and phenotypes, and potential drug targets. In evolutionary biology, sequencing is used to study how different organisms are related and how they evolved. In medicine, DNA sequencing is increasingly used to diagnose and treat rare diseases, identify genetic diseases, improve disease management, and provide reproductive counseling. It also helps determine specific bacteria for more precise antibiotic treatments, reducing the risk of antimicrobial resistance. In virology, sequencing is one of the main tools to identify and study viruses, and it can estimate when a viral outbreak began using a molecular clock technique. In forensic investigation, DNA sequencing is used along with DNA profiling methods for identification and paternity testing.
Did You Know?
- The first DNA sequences were obtained in the early 1970s using two-dimensional chromatography.
- There are over 10 million viral sequences in GenBank, though the exact number of unique sequences is not accurately represented.
The Technical Gauntlet of Working with a Single Cell
A typical human cell carries roughly 6.6 billion base pairs of DNA alongside hundreds of millions of mRNA bases, yet isolating just one cell means handling picogram quantities of nucleic acid. That vanishingly small starting material makes every stage of the workflow—extraction, amplification, library construction, sequencing, and downstream bioinformatic analysis—far more fragile than bulk sequencing of millions of pooled cells. Degradation, sample loss, and contamination loom large because there is simply no excess material to absorb a mistake. To compensate, protocols demand heavy amplification, which in turn introduces uneven coverage, background noise, and inaccurate quantification. Despite these obstacles, the payoff is extraordinary: single-cell resolution can expose heterogeneous samples, rare cell types, lineage relationships, somatic mosaicism, unculturable microbes, and the stepwise evolution of disease—problems that remain essentially invisible when cells are measured as a bulk mixture.
Amplification Strategies and Their Trade-offs
Multiple displacement amplification and MALBAC are the two dominant routes for boosting femtogram quantities of single-cell DNA into microgram-scale sequencing libraries. MALBAC also starts with isothermal amplification but appends a common sequence to its primers, prompting self-ligation into loops that halt further extension; a subsequent PCR cycle then amplifies the looped fragments, avoiding the highly branched networks typical of MDA. In practice, MDA delivers superior overall genome coverage and is better suited for SNP detection, while MALBAC produces more even coverage and excels at copy-number variant calls. Encapsulating cells in microfluidic droplets has further reduced bias and boosted throughput for MDA, though MALBAC's chemistry has not shown comparable gains from nanoliter-scale encapsulation.
Illuminating Heterogeneity in Cancer, Development, and Microbiomes
In oncology, pooling tumor cells masks the mutations carried by tiny subpopulations; single-cell DNA sequencing peels that layer away, revealing intra-tumor genetic heterogeneity and the role of somatic mosaicism in both disease progression and treatment response. In developmental biology, sequencing the RNA transcripts of individual cells exposes the existence and behavior of distinct cell types that bulk measurements blur into an average. In microbial ecology, a population of the same species can look genetically clonal at the bulk level, yet single-cell RNA or epigenetic profiling uncovers cell-to-cell variability that may underpin rapid adaptation to shifting environments. Perhaps most strikingly, single-cell DNA sequencing has opened the door to uncultivated prokaryotes in complex microbiomes: a genome recovered from one unicellular organism is termed a single amplified genome. Although individual SAGs suffer from low completeness and bias, computational tools such as SPAdes, IDBA-UD, Cortex, and HyDA can assemble near-complete genomes from composite SAGs, and the resulting data may eventually guide culturing strategies for organisms that have never grown in a laboratory.
Strand-seq and the Full Spectrum of Structural Variation
While MDA and MALBAC excel at detecting point mutations and copy-number changes, a different challenge—discovering large-scale structural rearrangements in a single cell—demanded a purpose-built method. Strand-seq, also called single-cell DNA template strand sequencing, addresses this by exploiting the orientation of reads inherited from each parental strand. The technique uses a tri-channel processing framework that jointly models read orientation, read depth, and haplotype phase within a single cell. This integrated approach enables the detection of the full spectrum of somatic structural variation classes at sizes of 200 kilobases or greater, a resolution that conventional single-cell amplification pipelines struggle to achieve. By preserving strand-specific information rather than collapsing it during random amplification, Strand-seq sidesteps many of the coverage-unevenness artifacts that plague MDA and MALBAC, making it a uniquely powerful tool for mapping the structural architecture of individual genomes and, by extension, the clonal evolution of cells within a tissue.
Gallery






Frequently Asked Questions
Who is DNA sequencing?
DNA sequencing is the laboratory process of reading the exact order of nucleotides—adenine, thymine, cytosine, and guanine—within a strand of DNA. It encompasses every technique, from classic gel-based approaches to modern high-throughput platforms, that reveals this four-letter genetic code.
What's DNA sequencing's origin story?
The first DNA sequences were obtained in the early 1970s, when researchers relied on two-dimensional chromatography to work out nucleotide order. This slow, labor-intensive method laid the groundwork for the rapid, high-throughput technologies that followed.
What are DNA sequencing's powers and range?
It can read genetic material stretching back over a million years, as demonstrated by the successful sequencing of mammoth DNA. It also maintains a vast archive of viral genomes, with GenBank holding more than 2.3 million viral sequences.
Why is DNA sequencing important to the canon?
It serves as a key technology across biology, medicine, forensics, and anthropology, enabling everything from disease diagnosis to identifying ancient human remains. The advent of rapid sequencing methods has dramatically accelerated research and discovery in all of these fields.
How does DNA sequencing's story end?
Rather than having a fixed ending, DNA sequencing continues to evolve, with each new generation of faster and cheaper methods opening fresh frontiers in genomics. Its ongoing expansion into clinical diagnostics and ecological genomics suggests the narrative is still being written.
Elsewhere in the Microbiology universe
Spotted an error? Know more?
This is a living reference — every entry is fact-audited, and reader corrections feed straight into our audit queue. Suggest an edit · See this site's audit record
