Glottolog: A Comprehensive Bibliographic Database of the World's Languages
In the vast field of linguistics, having access to accurate, verified, and organized information about the world's diverse languages is essential. Glottolog serves as a premier open-access online bibliographic database designed to meet this need. Produced by the Max Planck Institute of Geoanthropology in Germany, this resource provides researchers with a detailed catalogue of languages, language families, and extensive bibliographic materials.
Beyond simply listing names, Glottolog offers deep insights into the linguistic landscape by providing grammars, articles, and dictionaries that describe individual languages. It also maintains up-to-date language affiliations, curated by expert linguists to ensure scientific accuracy.
Key Facts
- Producer: Max Planck Institute of Geoanthropology (Germany).
- Access: Free and open-access.
- Core Function: Provides a catalogue of world languages and a comprehensive bibliography.
- Classification Style: Conservative regarding language families; permissive regarding language isolates.
- Data Standards: Uses ISO 639-3 codes and unique Glottocodes.
- Latest Version: 5.3 (released March 2026 under Creative Commons Attribution 4.0 International License).
The Origins and Purpose of Glottolog
The Glottolog/Langdoc project was established in 2011 by Sebastian Nordhoff and Harald Hammarström. The project was born out of a necessity for a more comprehensive language bibliography, addressing gaps found in other existing resources like Ethnologue. Initially developed at the Max Planck Institute for Evolutionary Anthropology in Leipzig, the project transitioned to the Max Planck Institute of Geoanthropology in Jena between 2015 and 2020.
Glottolog distinguishes itself through several rigorous methodological approaches:
- Verification: It only includes languages that editors have confirmed to exist and be distinct. Varieties that lack such confirmation are tagged as "spurious" or "unattested."
- Strict Classification: It only classifies languages into families that have been demonstrated to be valid groupings through specialized research.
- Bibliographic Depth: It provides extensive information, particularly for underdescribed or lesser-known languages.
- Standardized Coding: Language names in bibliographic entries are identified via ISO 639-3 or specific Glottocodes to facilitate easy cross-referencing with other databases.
Language Classification and Diversity
Glottolog employs a specific philosophy regarding language families—groups of languages related through descent from a common ancestor. While the database is conservative in establishing membership in large families, it is more permissive in classifying unclassified languages as isolates (languages with no demonstrable genealogical relationship to others).
In Edition 4.8, the database identified 421 spoken language families and isolates. The following table provides a snapshot of some of the most prominent groups recorded in the database:
| Name | Region | Number of Languages |
|---|---|---|
| Atlantic-Congo | Africa | 1,410 |
| Austronesian | Africa, Eurasia, Oceania, South America | 1,272 |
| Indo-European | Africa, Australia, Eurasia, North America, Oceania, South America | 585 |
| Sino-Tibetan | Eurasia | 506 |
| Afro-Asiatic | Africa, Eurasia | 382 |
| Nuclear Trans New Guinea | Oceania | 317 |
| Pama-Nyungan | Australia, Oceania | 250 |
| Otomanguean | North America | 181 |
| Austroasiatic | Eurasia | 158 |
| Tai-Kadai | Eurasia | 96 |
Beyond genealogical families, Glottolog also categorizes various other linguistic forms using non-genealogical classifications. These include pidgins (simplified languages used for communication between different groups), mixed languages, artificial languages, and sign languages.
Non-Genealogical Classifications
To ensure a complete linguistic picture, Glottolog tracks several categories that do not follow traditional family trees:
- Sign Languages: 223 entries, including those grouped as village sign languages.
- Pidgins: 84 languages.
- Unclassifiable Attested Languages: 121 entries.
- Artificial Languages: 31 entries.
- Mixed Languages: 9 entries.
- Speech Registers: 15 entries.
- Unattested Languages: 68 entries.
- Bookkeeping: 390 entries (including spurious languages and retired ISO entries).
Frequently Asked Questions
How does Glottolog differ from Ethnologue?
Glottolog focuses heavily on providing a comprehensive bibliography, especially for lesser-known languages. It is more conservative in its classification of language families and only includes languages that have been verified as distinct and existing, whereas it uses specific tags for unconfirmed varieties.
What is a Glottocode?
A Glottocode is a unique identifier used within the Glottolog database to identify specific languages, ensuring precision in bibliographic entries and facilitating links to other databases like ISO and Ethnologue.
Does Glottolog provide demographic or ethnographic data?
No. Glottolog focuses on linguistic and bibliographic data. It provides a single point-location on a map representing the geographic center of a language, but it does not include ethnographic or demographic information.
How are Creoles classified in the database?
In Glottolog, Creoles are classified alongside the specific language that provided their basic lexicon.
Is the data in Glottolog free to use?
Yes, Glottolog is an open-access resource. The latest version is released under the Creative Commons Attribution 4.0 International License.