GPT-1
OpenAI showed that a transformer pre-trained to predict the next word on unannotated text beats specialised models after fine-tuning.
Why it matters
The scheme of training on all text then fine-tuning on a task worked for language for the first time, and became the only scheme for the next decade.
The model had 117 million parameters and trained on a corpus of books. What matters is not the architecture but the order of operations: next-word prediction on an enormous unannotated corpus first, then brief fine-tuning for a task. Four months later BERT showed the same for bidirectional encoding, and the argument over which scheme was better ran until GPT-3.