DNA sequencingnucleotide orderSanger sequencinghigh-throughput sequencinggenomics

DNA Sequencing: Technologies, History, and Modern Applications

DNA Sequencing: Technologies, History, and Modern Applications DNA sequencing is the fundamental process of determining the nucleic acid sequence, which refers to the specific order of nu...

DNA Sequencing: Technologies, History, and Modern Applications

DNA sequencing is the fundamental process of determining the nucleic acid sequence, which refers to the specific order of nucleotides within a DNA molecule. This process involves identifying the arrangement of the four canonical bases: adenine (A), thymine (T), cytosine (C), and guanine (G). By mapping these sequences, scientists can unlock the biological blueprints of living organisms.

The ability to sequence DNA rapidly has revolutionized biological and medical research. From diagnosing diseases like cancer to advancing biotechnology and forensic biology, the precision of sequencing allows for highly individualized medical care and the systematic cataloging of the world's biodiversity.

An example of the results of automated chain-termination DNA sequencing
An example of the results of automated chain-termination DNA sequencing

Key Facts

  • DNA sequencing determines the order of the four bases: adenine, thymine, cytosine, and guanine.
  • Modern sequencing is used in medical diagnosis, forensics, virology, and evolutionary biology.
  • Technological shifts have moved from laborious chromatography to rapid, high-throughput automated methods.
  • Sequencing enables personalized medicine by comparing healthy and mutated DNA.

The Evolution of Sequencing Technology

The journey of genomic discovery began in the early 1970s. Initially, academic researchers relied on laborious methods based on two-dimensional chromatography to obtain the first DNA sequences. These early processes were time-consuming and difficult to scale.

Frederick Sanger, a pioneer of sequencing. Sanger is one of the few scientists who was awarded two Nobel prizes, one for the sequencing of proteins, and the other for the sequencing of DNA.
Frederick Sanger, a pioneer of sequencing. Sanger is one of the few scientists who was awarded two Nobel prizes, one for the sequencing of proteins, and the other for the sequencing of DNA.

A major turning point occurred with the development of chain-termination methods, most notably pioneered by Frederick Sanger. This advancement, along with the introduction of fluorescence-based sequencing, made the process significantly easier and orders of magnitude faster than previous techniques.

History of sequencing technology [64]
History of sequencing technology [64]

Early Methods and the Sanger Legacy

One of the most influential early techniques was Sanger sequencing, which utilizes chain-terminating inhibitors to stop DNA synthesis at specific points, allowing the sequence to be read. While newer methods have emerged, the principles established during this era remain foundational to the field.

The 5,386 bp genome of bacteriophage φX174. Each coloured block represents a gene.
The 5,386 bp genome of bacteriophage φX174. Each coloured block represents a gene.

Modern Sequencing Methodologies

Today, sequencing is categorized into several distinct approaches, primarily distinguished by the length of the DNA fragments they can read and the speed at which they operate.

Short-Read Sequencing

Short-read sequencing is highly accurate and widely used for high-throughput applications. Technologies like Illumina sequencing (sequencing by synthesis) dominate this space. These methods involve breaking genomic DNA into smaller pieces, cloning them, and then sequencing the overlapping regions to reconstruct the full genome.

Genomic DNA is fragmented into random pieces and cloned as a bacterial library. DNA from individual bacterial clones is sequenced and the sequence is assembled by using overlapping DNA regions.
Genomic DNA is fragmented into random pieces and cloned as a bacterial library. DNA from individual bacterial clones is sequenced and the sequence is assembled by using overlapping DNA regions.
Multiple, fragmented sequence reads must be assembled together on the basis of their overlapping areas.
Multiple, fragmented sequence reads must be assembled together on the basis of their overlapping areas.

Other short-read technologies include Ion Torrent (semiconductor sequencing) and SOLiD (sequencing by ligation). While these methods offer high throughput, they typically produce shorter fragments compared to newer technologies.

Library preparation for the SOLiD platform
Library preparation for the SOLiD platform
Two-base encoding scheme. In two-base encoding, each unique pair of bases on the 3' end of the probe is assigned one out of four possible colors. For example, "AA" is assigned to blue, "AC" is assigned to green, and so on for all 16 unique pairs. During sequencing, each base in the template is sequenced twice, and the resulting data are decoded according to this scheme.
Two-base encoding scheme. In two-base encoding, each unique pair of bases on the 3' end of the probe is assigned one out of four possible colors. For example, "AA" is assigned to blue, "AC" is assigned to green, and so on for all 16 unique pairs. During sequencing, each base in the template is sequenced twice, and the resulting data are decoded according to this scheme.

Long-Read Sequencing

To overcome the limitations of short reads, long-read sequencing methods have been developed. These technologies, such as Pacific Biosciences (SMRT) and Nanopore sequencing, can read much longer continuous stretches of DNA. This is particularly useful for assembling complex genomes and identifying structural variations.

Sequencing of the TAGGCT template with IonTorrent, PacBioRS and GridION
Sequencing of the TAGGCT template with IonTorrent, PacBioRS and GridION

Comparison of Sequencing Technologies

The following table provides a technical comparison of various sequencing methods currently used in research and clinical settings.

Comparison of Major Sequencing Methods
Method Read Length Accuracy Cost per 1B Bases (USD)
Sanger (Chain Termination) 400–900 bp 99.9% $2,400,000
Illumina (Sequencing by Synthesis) 50–600 bp 99.9% $5–$150
Ion Torrent (Semiconductor) Up to 600 bp 99.6% $66.8–$950
Pacific Biosciences (SMRT) 30,000 bp (N50) 87% (raw) $7.2–$43.3
Nanopore Variable 92–97% $7–$100

Note: Costs and performance metrics can vary significantly based on the specific instrument and run parameters used.

An Illumina HiSeq 2500 sequencer
An Illumina HiSeq 2500 sequencer
Illumina NovaSeq 6000 flow cell
Illumina NovaSeq 6000 flow cell
An Illumina MiSeq sequencer
An Illumina MiSeq sequencer
A BGI MGISEQ-2000RS sequencer
A BGI MGISEQ-2000RS sequencer

Applications of DNA Sequencing

The impact of sequencing spans across multiple scientific disciplines:

  • Medicine: Comparing healthy and mutated DNA to diagnose cancers and guide personalized patient treatments.
  • Molecular Biology: Studying the fundamental processes of life at a genetic level.
  • Evolutionary Biology: Understanding the relationships between different species through genomic comparison.
  • Forensics: Using DNA profiles to assist in criminal investigations.
  • Virology: Identifying and tracking the mutations of viruses.
  • Metagenomics: Analyzing genetic material recovered directly from environmental samples.
Total cost of sequencing a human genome over time as calculated by the NHGRI
Total cost of sequencing a human genome over time as calculated by the NHGRI

Frequently Asked Questions

What are the four bases of DNA?

The four canonical bases that make up the DNA sequence are adenine (A), thymine (T), cytosine (C), and guanine (G).

What is the difference between short-read and long-read sequencing?

Short-read sequencing produces many small, highly accurate fragments of DNA, which are then assembled like a puzzle. Long-read sequencing produces much longer continuous sequences, making it easier to map complex or repetitive regions of a genome.

How has the cost of sequencing changed over time?

The cost of sequencing a human genome has decreased dramatically over time, moving from millions of dollars to a range that is accessible for large-scale research and clinical applications.

Why is DNA sequencing important for medicine?

Sequencing allows doctors to identify specific genetic mutations associated with diseases like cancer. This information can be used to tailor treatments to a patient's unique genetic makeup, a field known as personalized medicine.

What is high-throughput sequencing?

High-throughput sequencing, also known as next-generation sequencing (NGS), refers to modern technologies that can sequence millions of DNA fragments simultaneously, allowing for massive amounts of data to be generated in a single run.