GPT-3
A model with 175 billion parameters performed tasks from a few examples in the prompt, with no weight updates at all.
Why it matters
A capability nobody programmed appeared: the model learns from the statement of the task rather than from a gradient step.
A hundredfold jump over GPT-2 produced not merely better quality but qualitatively different behaviour. In-context learning meant a new task needed neither data nor retraining, only a description in the prompt. That changed how a model is used: the prompt replaced fine-tuning. Together with the scaling laws of the same year, GPT-3 fixed the direction toward size.