NCBIbioinformaticsGenBankPubMedBLAST

National Center for Biotechnology Information (NCBI): The Hub of Biomedical Data

National Center for Biotechnology Information (NCBI): The Hub of Biomedical Data The National Center for Biotechnology Information (NCBI) serves as a cornerstone of modern biomedical rese...

National Center for Biotechnology Information (NCBI): The Hub of Biomedical Data

The National Center for Biotechnology Information (NCBI) serves as a cornerstone of modern biomedical research. Established as a part of the National Library of Medicine (NLM), which is a branch of the National Institutes of Health (NIH), the NCBI is a U.S. government-funded institution dedicated to the storage, standardization, and distribution of biotechnology information.

Based in Bethesda, Maryland, the NCBI provides the global scientific community with essential bioinformatics tools—software used to analyze biological data—and a vast array of databases that integrate genetic, protein, and chemical information.

The Lister Hill Center, the NIH campus building that has hosted the NCBI since its creation in 1988
The Lister Hill Center, the NIH campus building that has hosted the NCBI since its creation in 1988

Key Facts

  • Founded: 1988, through legislation sponsored by Congressman Claude Pepper.
  • Parent Organization: United States National Library of Medicine (NLM).
  • Primary Purpose: To standardize biotechnology databases and provide analysis tools for researchers.
  • Core Tools: BLAST for sequence alignment and Entrez for cross-database searching.
  • Major Repositories: GenBank (DNA), PubMed (literature), and PubChem (chemicals).
  • Current Director: Stephen Sherry (since September 26, 2022).

History and Evolution

Proposal and Establishment

In the mid-1980s, the NLM identified a critical problem: biotechnology databases created by different organizations lacked consistent naming schemes and connectivity. This fragmentation made it difficult for researchers to integrate data from multiple sources. To solve this, the NLM proposed the creation of the NCBI to lead standardization efforts and act as a central repository for biotechnology information.

Following several bills introduced by Congressman Claude Pepper, the NCBI was officially established on November 4, 1988.

Technological Milestones

The NCBI has consistently evolved to meet the needs of genomic science. In 1990, a team including mathematician Stephen Altschul developed BLAST, a revolutionary tool for comparing biological sequences. By 1992, the NCBI assumed responsibility for GenBank, the primary database for DNA and amino acid sequences, initially providing access via email.

The launch of the NCBI website in 1994 transformed accessibility, bringing BLAST, GenBank, and the Entrez retrieval system to the web. Later milestones include hosting the human genome mapped by the Human Genome Project in 2000 and the creation of PubMed Central (PMC), a free archive of biomedical literature. In 2008, a U.S. Congress mandate required NIH-funded research papers to be made public on PubMed within 12 months of publication.

Homepage of the NCBI website showing access to major databases and tools, including the "Popular Resources" panel and navigation menu
Homepage of the NCBI website showing access to major databases and tools, including the "Popular Resources" panel and navigation menu

Core NCBI Resources

The NCBI ecosystem is divided into literature resources, analysis tools, and molecular databases, all interconnected to streamline research.

Literature Resources

  • PubMed: A primary database for searching citations and abstracts of biomedical research.
  • PubMed Central (PMC): A free, full-text archive of journal literature used for research and text mining.
  • Bookshelf: A digital repository of biomedical books, clinical manuals, and monographs contributed by government agencies and academic publishers.
NCBI Bookshelf homepage interface showing access to biomedical literature, search tools, and integrated database navigation
NCBI Bookshelf homepage interface showing access to biomedical literature, search tools, and integrated database navigation
Browse titles interface in the NCBI Bookshelf, displaying available publications and filtering options by category, publisher, and date
Browse titles interface in the NCBI Bookshelf, displaying available publications and filtering options by category, publisher, and date

Analysis Tools: BLAST

The Basic Local Alignment Search Tool (BLAST) is an algorithm used to find regions of similarity between a query sequence and sequences in a database. Because full sequence comparison is computationally expensive, BLAST uses heuristic methods—shortcuts that prioritize speed while maintaining accuracy.

The process involves identifying short matching sequences, extending them to find initial alignments, and finally performing a "traceback" to include insertions and deletions for a detailed final alignment. Different versions of BLAST exist for different needs:

  • blastn: Nucleotide query vs. nucleotide database.
  • blastp: Protein query vs. protein database.
  • blastx: Translated nucleotide query vs. protein database.
  • tblastn: Protein query vs. translated nucleotide database.
Example of the blastp interface showing fields for entering a query sequence, selecting a database, and adjusting search parameters
Example of the blastp interface showing fields for entering a query sequence, selecting a database, and adjusting search parameters

Molecular and Genomic Databases

The NCBI manages a diverse array of specialized databases:

  • GenBank: An annotated collection of all publicly available DNA sequences, coordinated with international partners like EMBL and DDBJ.
  • Gene: A directory providing organized information about genes across species, using unique GeneIDs.
  • Protein: A collection of protein sequences derived from sources like RefSeq and UniProtKB/SWISS-Prot.
  • PubChem: A resource for chemical molecules and their biological activities, consisting of three parts: Substance, Compound, and BioAssay.
  • dbSNP: A database focusing on single-nucleotide variations and small-scale insertions/deletions in humans.
  • Sequence Read Archive (SRA): Stores raw sequencing data from next-generation platforms (e.g., Illumina, Pacific Biosciences).
Diagram exemplifying how the different databases that compose PubChem interact with each other and with the user, as depicted by Rosania et al., 2007[17]
Diagram exemplifying how the different databases that compose PubChem interact with each other and with the user, as depicted by Rosania et al., 2007[17]

The Entrez System

Entrez is the global query cross-database search system that powers the NCBI. Rather than searching each database individually, Entrez integrates data from various sources—including PubMed, GenBank, and Taxonomy—into a uniform information model. This allows researchers to efficiently retrieve related references, sequences, and structures through a single interface.

Summary of Major NCBI Databases

Overview of Primary NCBI Databases and Their Functions
Database Primary Content Key Use Case
PubMed Biomedical citations/abstracts Literature review and research discovery
GenBank Public DNA sequences Genetic sequence retrieval and storage
PubChem Chemical molecules and bioassays Chemical property and activity research
Gene Gene characterization Mapping gene function and homology
Protein Amino acid sequences Protein structure and function analysis
ClinVar Genomic variations Linking genetic variants to human health

Frequently Asked Questions

What is the difference between PubMed and PubMed Central?

PubMed is a database of citations and abstracts for biomedical literature, whereas PubMed Central (PMC) is a free digital archive that provides the full-text versions of those articles.

How does BLAST speed up sequence searching?

BLAST uses heuristic methods to find short matches first and only performs computationally expensive detailed alignments on sequences that meet a specific score threshold.

What is the purpose of the Entrez system?

Entrez acts as a unified search and retrieval system that allows users to query multiple NCBI databases simultaneously using a single search interface.

Who provides the content for the NCBI Bookshelf?

Content is contributed by scientific publishers, academic institutions, and government agencies, ensuring the materials are peer-reviewed and scientifically accurate.

What is the role of GenBank in the global community?

GenBank serves as the NIH's genetic sequence database and coordinates with the European Molecular Biology Laboratory (EMBL) and the DNA Data Bank of Japan (DDBJ) to maintain a global collection of DNA sequences.