Lancaster-Oslo/Bergen (LOB) Corpus: A Pillar of British English Linguistics

Lancaster-Oslo/Bergen (LOB) Corpus: A Pillar of British English Linguistics

In the field of linguistics, a corpus—a large, structured collection of texts—serves as the essential raw material for analyzing how language is actually used. One of the most significant milestones in this discipline is the Lancaster-Oslo/Bergen (LOB) Corpus, a comprehensive one-million-word collection of British English texts.

Developed in the 1970s, the LOB Corpus was created through a strategic collaboration between the University of Lancaster, the University of Oslo, and the Norwegian Computing Centre for the Humanities in Bergen. Its primary purpose was to provide a British English counterpart to the Brown Corpus, which had been compiled by Henry Kučera and W. Nelson Francis in the 1960s to represent American English.

Key Facts

  • Total Size: Approximately one million words.
  • Origin: Compiled in the 1970s by the University of Lancaster, University of Oslo, and the Norwegian Computing Centre for the Humanities.
  • Primary Goal: To mirror the Brown Corpus (American English) using British English texts.
  • Source Material: Documents published in the UK in 1961 by British authors.
  • Structure: 500 samples, each containing roughly 2,000 words.
  • Key Figures: Chief compilers Geoffrey Leech and Stig Johansson.
  • Enhancements: The corpus is fully tagged for part-of-speech categories.

Design and Composition

To ensure a scientifically valid comparison between British and American English, the LOB Corpus was designed to match the Brown Corpus as closely as possible in terms of size and genre. By utilizing texts from 1961, the compilers created a snapshot of the English language during a specific era.

The collection is divided into 500 distinct samples. Each sample consists of about 2,000 words, distributed across a wide array of categories ranging from journalistic reportage to scientific writings and various forms of fiction.

Label Text Category Brown Corpus (US) LOB Corpus (UK)
A Press: reportage 44 44
B Press: editorial 27 27
C Press: reviews 17 17
D Religion 17 17
E Skills, trades and hobbies 36 38
F Popular lore 48 44
G Belles lettres, biography, essays 75 77
H Miscellaneous (documents, reports, etc.) 30 30
J Learned and scientific writings 80 80
K General fiction 29 29
L Mystery and detective fiction 24 24
M Science fiction 6 6
N Adventure and western fiction 29 29
P Romance and love story 29 29
R Humour 9 9
Total - 500 500

Technical Enhancements: Part-of-Speech Tagging

Beyond the mere collection of text, the LOB Corpus has been tagged. In linguistics, tagging refers to the process of assigning a part-of-speech (POS) category to every single word in the corpus. This means each word is labeled as a noun, verb, adjective, or other grammatical category, allowing researchers to perform complex computational analyses of grammatical patterns and syntactic structures.

Frequently Asked Questions

What is the main difference between the LOB and Brown corpora?

The primary difference is the regional variety of English they represent: the Brown Corpus focuses on American English, while the LOB Corpus focuses on British English.

Who were the primary creators of the LOB Corpus?

The chief compilers were Geoffrey Leech from Lancaster University and Stig Johansson from the University of Oslo.

When were the texts in the LOB Corpus published?

The corpus consists of documents published in the United Kingdom in 1961 by British authors.

How is the LOB Corpus structured?

It contains 500 samples, with each sample comprising approximately 2,000 words, totaling roughly one million words across various genres.

What does it mean that the LOB Corpus is "tagged"?

Tagging means that every word in the collection has been assigned a specific part-of-speech category, facilitating deeper linguistic and grammatical research.