Data outweighs the algorithm
Banko and Brill trained four different algorithms on corpora from one million to one billion words and showed that they converge at large scale.
Why it matters
The simplest algorithm on a billion words beat the best one on a million, and that moved the field's effort from methods to data.
The curves for all four methods kept rising and had not flattened even at a billion words. The authors stated the conclusion carefully: at available scales the choice of algorithm matters less than had been assumed. Eight years later Halevy, Norvig and Pereira expanded the same thesis into the unreasonable effectiveness of data. It is an ancestor of what was later named scaling laws.