ISO 639-2: The Standard for Alpha-3 Language Codes
In the world of global data exchange and linguistics, precisely identifying a language is critical. ISO 639-2:1998 is the international standard that provides a systematic way to represent language names using three-letter codes, known as Alpha-3 codes. With 487 entries, this standard ensures that libraries, software developers, and linguists can communicate language identities without ambiguity.
The administration of this standard is handled by the US Library of Congress, acting as the registration authority (ISO 639-2/RA). The Library of Congress reviews proposed changes and collaborates with the ISO 639-RA Joint Advisory Committee to maintain the accuracy of the code tables.
[ไม่มีภาพประกอบ]
Key Facts
- Standard Name: ISO 639-2:1998 (Part 2 of the ISO 639 standard).
- Code Format: Three-letter (Alpha-3) identifiers.
- Total Entries: 487 language codes.
- Registration Authority: The US Library of Congress.
- Primary Purpose: To provide a more expansive coding system than the two-letter ISO 639-1 standard.
History and Evolution of ISO 639 Standards
Development of ISO 639-2 began in 1989. The primary driver for its creation was the limitation of the ISO 639-1 standard; because ISO 639-1 uses only two-letter codes, it lacked the capacity to accommodate the vast number of known languages. ISO 639-2 was officially released in 1998 to solve this scalability issue.
Over time, the landscape of language coding has evolved. ISO 639-2 has been largely superseded by ISO 639-3 (released in 2007), which offers comprehensive coverage of all individual languages. While ISO 639-3 incorporates the individual language codes from ISO 639-2, it does not include collective languages (groups of related languages). Instead, most collective languages are now managed under ISO 639-5.
B Codes vs. T Codes
A unique characteristic of ISO 639-2 is that twenty languages are assigned two different three-letter codes. These are categorized as follows:
- Bibliographic codes (ISO 639-2/B): Derived from the English name of the language. These were maintained as legacy features.
- Terminological codes (ISO 639-2/T): Derived from the native name of the language and typically align more closely with the ISO 639-1 two-letter codes.
In modern application, T codes are generally preferred, and ISO 639-3 exclusively adopts the ISO 639-2/T versions. It is worth noting that there were originally 22 B codes, but scc and scr have since been deprecated.
Scopes and Classifications
ISO 639-2 codes are not one-size-fits-all; they are categorized by their scope of denotation, which defines the type of linguistic entity being identified.
Language Types
Individual languages are further classified into five distinct types:
- Living languages
- Extinct languages
- Ancient languages
- Historic languages
- Constructed languages
Broad Categories
Beyond individual languages, the standard includes codes for macrolanguages, dialects, and collections of languages. Some codes are also reserved for local use or specific special situations.
Collections of Languages
Collective language codes represent groups of related languages rather than a single specific tongue. These are excluded from ISO 639-3. These groups are divided into remainder groups (which exclude languages that already have their own individual codes) and inclusive groups.
| Code | Language Group |
|---|---|
| afa | Afro-Asiatic languages |
| cel | Celtic languages |
| gem | Germanic languages |
| ine | Indo-European languages |
| sem | Semitic languages |
| sla | Slavic languages |
| sgn | Sign languages |
Reserved and Special Codes
Local Use
The range from qaa to qtz is reserved for local use. These codes are not official ISO 639-2 or 639-3 assignments and are used privately for languages not yet standardized. For example, Microsoft Windows uses qps for pseudo-locales to test software localization.
Special Situations
Four generic codes are used to handle edge cases in linguistic data:
- mis: Uncoded languages (originally "miscellaneous").
- mul: Multiple languages, used when several languages are present and listing each is impractical.
- und: Undetermined, used when a language must be indicated but cannot be identified.
- zxx: No linguistic content (e.g., animal sounds), added on January 11, 2006.
Frequently Asked Questions
What is the difference between ISO 639-1 and ISO 639-2?
ISO 639-1 uses two-letter codes and is limited in the number of languages it can represent. ISO 639-2 uses three-letter (Alpha-3) codes to accommodate a much larger number of languages.
What are B and T codes in ISO 639-2?
B codes (Bibliographic) are based on the English name of a language, while T codes (Terminological) are based on the native name. T codes are generally preferred in modern usage.
Is ISO 639-2 still the primary standard?
For individual languages, ISO 639-2 has been largely superseded by ISO 639-3, which provides more comprehensive coverage. However, ISO 639-2 is still relevant for collective language codes.
What does the code 'zxx' represent?
The code 'zxx' is used to indicate that there is no linguistic content, such as in the case of animal sounds.
Who manages the ISO 639-2 registration?
The US Library of Congress serves as the registration authority (ISO 639-2/RA), reviewing changes and maintaining the code tables.