Back to timeline

Research · December 1992

Word classes from mutual information

In December 1992 five IBM researchers led by Peter Brown described in Computational Linguistics how to divide a vocabulary of 260,741 words into 1,000 classes purely from which words stand next to each other in almost 366 million words of text, and how to build a language model on classes instead of single words.

Why it matters

Word similarity was derived from text alone, without a dictionary or annotation, and the classes came out both grammatical and semantic. Bengio's neural language model of 2003 starts from exactly this idea, replaces a word's discrete class with a continuous vector, and compares itself with class-based models.

The algorithm is greedy: every word starts as its own class, and the pair of classes whose merger loses the least average mutual information between adjacent classes is merged; for a large vocabulary words are added in order of frequency. Perplexity of the Brown corpus (1,014,312 words, not used in training): 244 for the word trigram model, 271 for the class model, 236 for the two mixed. Classes cut the share of unseen trigrams from 14.7% to 3.8%. The same paper finds "sticky" word pairs and semantic classes from co-occurrence at a distance. What the record does not claim. The gain in perplexity is small: the class model alone is worse than the word model, and only the mixture helps.

Event record

Event date
December 1992
Timeline date
Event date
Verification
Sources gathered automatically · September 26, 2026
Lines
ID
evt-0842

Computational Linguistics, volume 18, issue 4, December 1992.

Sources

Related events