Chen and Goodman's comparison of smoothing
In August 1998 Stanley Chen and Joshua Goodman issued a Harvard technical report comparing the main smoothing methods for n-gram language models on the Brown, North American Business news, Switchboard and Broadcast News corpora across training set sizes. Their own modified variant of Kneser-Ney won every comparison.
Why it matters
Choosing a smoothing method stopped being a matter of habit: the report showed what each method's advantage depends on and supplied one that won everywhere. N-grams with modified Kneser-Ney smoothing became the baseline Bengio's neural language model measured itself against in 2003.
The methods of Jelinek-Mercer, Katz, Bell-Cleary-Witten, Ney-Essen-Kneser and Kneser-Ney were compared by the cross-entropy of test text. The modification: three discounts instead of one for all non-zero counts, for n-grams seen once, twice, and three or more times. In Broadcast News speech recognition each bit of cross-entropy corresponded to about 5.4% absolute word error, and the gap between the best and a mediocre smoothing method, 0.2 bits or more, to about 1%. What the record does not claim. Bengio cites the 1999 journal version of this work in Computer Speech and Language, not the 1998 report; the record is dated by the report because that is what was read.