A neural probabilistic language model
Bengio and co-authors replaced n-gram tables with a network that learns a vector for each word and predicts the next word from those vectors.
Why it matters
Words became points in a space where nearness means similar meaning, the basis of all later language modelling.
The problem with n-grams is sparsity: most word combinations never occur in a corpus at all. A distributed representation solves it by generalisation, so a model that has seen a cat running gives a sensible probability to a dog running. Computationally the model was too expensive for its time. Word2vec in 2013 and ELMo are direct descendants of the idea.