National Center for Biotechnology Information (NCBI): The Hub of Biomedical Data
The National Center for Biotechnology Information (NCBI) serves as a cornerstone of modern biomedical research. Established as a part of the National Library of Medicine (NLM), which is a branch of the National Institutes of Health (NIH), the NCBI is a U.S. government-funded institution dedicated to the storage, standardization, and distribution of biotechnology information.
Based in Bethesda, Maryland, the NCBI provides the global scientific community with essential bioinformatics tools—software used to analyze biological data—and a vast array of databases that integrate genetic, protein, and chemical information.

Key Facts
- Founded: 1988, through legislation sponsored by Congressman Claude Pepper.
- Parent Organization: United States National Library of Medicine (NLM).
- Primary Purpose: To standardize biotechnology databases and provide analysis tools for researchers.
- Core Tools: BLAST for sequence alignment and Entrez for cross-database searching.
- Major Repositories: GenBank (DNA), PubMed (literature), and PubChem (chemicals).
- Current Director: Stephen Sherry (since September 26, 2022).
History and Evolution
Proposal and Establishment
In the mid-1980s, the NLM identified a critical problem: biotechnology databases created by different organizations lacked consistent naming schemes and connectivity. This fragmentation made it difficult for researchers to integrate data from multiple sources. To solve this, the NLM proposed the creation of the NCBI to lead standardization efforts and act as a central repository for biotechnology information.
Following several bills introduced by Congressman Claude Pepper, the NCBI was officially established on November 4, 1988.
Technological Milestones
The NCBI has consistently evolved to meet the needs of genomic science. In 1990, a team including mathematician Stephen Altschul developed BLAST, a revolutionary tool for comparing biological sequences. By 1992, the NCBI assumed responsibility for GenBank, the primary database for DNA and amino acid sequences, initially providing access via email.
The launch of the NCBI website in 1994 transformed accessibility, bringing BLAST, GenBank, and the Entrez retrieval system to the web. Later milestones include hosting the human genome mapped by the Human Genome Project in 2000 and the creation of PubMed Central (PMC), a free archive of biomedical literature. In 2008, a U.S. Congress mandate required NIH-funded research papers to be made public on PubMed within 12 months of publication.

Core NCBI Resources
The NCBI ecosystem is divided into literature resources, analysis tools, and molecular databases, all interconnected to streamline research.
Literature Resources
- PubMed: A primary database for searching citations and abstracts of biomedical research.
- PubMed Central (PMC): A free, full-text archive of journal literature used for research and text mining.
- Bookshelf: A digital repository of biomedical books, clinical manuals, and monographs contributed by government agencies and academic publishers.


Analysis Tools: BLAST
The Basic Local Alignment Search Tool (BLAST) is an algorithm used to find regions of similarity between a query sequence and sequences in a database. Because full sequence comparison is computationally expensive, BLAST uses heuristic methods—shortcuts that prioritize speed while maintaining accuracy.
The process involves identifying short matching sequences, extending them to find initial alignments, and finally performing a "traceback" to include insertions and deletions for a detailed final alignment. Different versions of BLAST exist for different needs:
- blastn: Nucleotide query vs. nucleotide database.
- blastp: Protein query vs. protein database.
- blastx: Translated nucleotide query vs. protein database.
- tblastn: Protein query vs. translated nucleotide database.

Molecular and Genomic Databases
The NCBI manages a diverse array of specialized databases:
- GenBank: An annotated collection of all publicly available DNA sequences, coordinated with international partners like EMBL and DDBJ.
- Gene: A directory providing organized information about genes across species, using unique GeneIDs.
- Protein: A collection of protein sequences derived from sources like RefSeq and UniProtKB/SWISS-Prot.
- PubChem: A resource for chemical molecules and their biological activities, consisting of three parts: Substance, Compound, and BioAssay.
- dbSNP: A database focusing on single-nucleotide variations and small-scale insertions/deletions in humans.
- Sequence Read Archive (SRA): Stores raw sequencing data from next-generation platforms (e.g., Illumina, Pacific Biosciences).
![Diagram exemplifying how the different databases that compose PubChem interact with each other and with the user, as depicted by Rosania et al., 2007[17]](/images/2f/37/2f37853f69e14c6f789f10d05cf49fcb14417d97d2e45e8a1e09b1cd63878758.jpg)
The Entrez System
Entrez is the global query cross-database search system that powers the NCBI. Rather than searching each database individually, Entrez integrates data from various sources—including PubMed, GenBank, and Taxonomy—into a uniform information model. This allows researchers to efficiently retrieve related references, sequences, and structures through a single interface.
Summary of Major NCBI Databases
| Database | Primary Content | Key Use Case |
|---|---|---|
| PubMed | Biomedical citations/abstracts | Literature review and research discovery |
| GenBank | Public DNA sequences | Genetic sequence retrieval and storage |
| PubChem | Chemical molecules and bioassays | Chemical property and activity research |
| Gene | Gene characterization | Mapping gene function and homology |
| Protein | Amino acid sequences | Protein structure and function analysis |
| ClinVar | Genomic variations | Linking genetic variants to human health |
Frequently Asked Questions
What is the difference between PubMed and PubMed Central?
PubMed is a database of citations and abstracts for biomedical literature, whereas PubMed Central (PMC) is a free digital archive that provides the full-text versions of those articles.
How does BLAST speed up sequence searching?
BLAST uses heuristic methods to find short matches first and only performs computationally expensive detailed alignments on sequences that meet a specific score threshold.
What is the purpose of the Entrez system?
Entrez acts as a unified search and retrieval system that allows users to query multiple NCBI databases simultaneously using a single search interface.
Who provides the content for the NCBI Bookshelf?
Content is contributed by scientific publishers, academic institutions, and government agencies, ensuring the materials are peer-reviewed and scientifically accurate.
What is the role of GenBank in the global community?
GenBank serves as the NIH's genetic sequence database and coordinates with the European Molecular Biology Laboratory (EMBL) and the DNA Data Bank of Japan (DDBJ) to maintain a global collection of DNA sequences.