Back to timeline

Research · February 9 – 12, 1988

Statistical part-of-speech tagging

At the ANLP conference in Austin in February 1988 Kenneth Church of Bell Laboratories presented a program that assigns parts of speech in unrestricted text by multiplying two probabilities estimated on the tagged Brown Corpus of about a million words: how often a word takes a part of speech, and how often that part follows the two before it. It was 95-99% correct, depending on the definition of correct.

Why it matters

Part-of-speech tagging, until then a matter of grammatical rules, came down to counting in a corpus and linear-time dynamic programming. Church's PARTS program provided the initial automatic tagging of the Penn Treebank, which annotators then corrected.

The lexical probability is the share of cases in which a word takes a part of speech in the Brown Corpus (I is a pronoun in 5,837 cases out of 5,838); the contextual one is a tag trigram frequency divided by the bigram frequency. Church estimates the uncertainty at about 0.25 bits per tag when the word is known and about 2 bits when only the two neighbouring tags are. A 400-word sample was judged 99.5% correct. The same approach found simple noun phrases, their probabilities estimated from 40,000 words (11,000 phrases). What the record does not claim. Church was not the first to apply n-grams to tagging: the paper itself names Mercer's experiments at IBM in 1982 and the Lancaster group that tagged the LOB corpus with 96.7% success, and says it developed independently of that work. The 95-99% figure is the author's estimate, not a measurement on held-out text.

Event record

Event date
February 9, 1988 – February 12, 1988
Timeline date
Event date
Verification
Sources gathered automatically · September 26, 2026
Lines
ID
evt-0839

The Second Conference on Applied Natural Language Processing (ANLP), Austin, Texas: the ACL Anthology gives the month, Crossref the days, 9-12 February 1988.

Sources

Related events

Records that link to this one