word2vec
Mikolov and co-authors simplified the neural language model enough that word vectors could be trained on billions of words in hours.
Why it matters
Word vectors went from a research curiosity to an everyday tool available to anyone with a laptop.
The key was removing the hidden layer and replacing the full softmax with approximations. The best-known result is arithmetic on vectors: king minus man plus woman gives queen. That made visible what the model learns is the structure of meaning, not only co-occurrence statistics. The toolkit spread instantly and was the first step toward the representations BERT was later built on.