The IBM translation models
The paper set out five word-alignment models in detail and a way to estimate their parameters from an unannotated bilingual corpus.
Why it matters
Machine translation acquired a reproducible recipe that held until neural translation arrived in 2014.
Models one through five grow steadily more complex: from a simple word correspondence to accounting for order and fertility, that is how many words one word maps to. Parameters are estimated by expectation maximisation. Word alignments from these models remained a basic step in the phrase-based systems of the 2000s and in training sets for neural systems.