Maximum entropy for language processing
In March 1996 Adam Berger and Stephen and Vincent Della Pietra described in Computational Linguistics how to build a probabilistic model of language on the principle of maximum entropy: of all the distributions consistent with chosen facts about the data, take the most uniform, and select the facts (features) automatically.
Why it matters
Any features of the context, such as neighbouring words, their forms or position, could now be combined in one model without assuming they were independent. Conditional random fields in 2001 grew out of a limit of exactly such local models.
The work was done at the IBM T. J. Watson Research Center with ARPA support; by publication Berger was at Columbia University and the Della Pietras at Renaissance Technologies. Parameters are estimated by Improved Iterative Scaling, a version of Darroch and Ratcliff's iterative scaling of 1972; features are added greedily by gain in likelihood. The examples come from IBM's Candide statistical translator (French to English) on the Canadian Hansard corpus of 3.6 million sentence pairs: translating the word in by context, cutting a French sentence into segments of at most 10 words that can be translated in turn, and reordering NOUN de NOUN phrases. In reordering, a model of 358 constraints trained on 10,000 examples was right on 80.4% of 71,555 test phrases, against 70.2% for the rule never to swap. What the record does not claim. The principle and the scaling algorithm are older than the paper; what is new in it is automatic feature selection and the application to language.