Back to timeline

Research · March 29, 2022

Chinchilla: compute-optimal training

On 29 March 2022 Jordan Hoffmann and colleagues at DeepMind posted the result of training over 400 language models from 70 million to over 16 billion parameters on 5 to 500 billion tokens: for compute-optimal training, model size and the number of tokens should be scaled equally. Chinchilla, 70 billion parameters trained on 1.4 trillion tokens with Gopher's budget (4 times more data than the 280-billion-parameter Gopher), outperformed Gopher, GPT-3, Jurassic-1 and Megatron-Turing NLG and reached 67.5 per cent on MMLU, a greater than 7 per cent improvement over Gopher.

Why it matters

The rule by which large models were sized was rewritten: by Kaplan and co-authors (2020) a tenfold budget meant a 5.5-fold larger model on only 1.8 times more tokens, and most large models had been trained on about 300 billion tokens. By this paper they were significantly undertrained. LLaMA in 2023 takes this recommendation as its point of departure.

Every doubling of model size should come with a doubling of training tokens. Being smaller than Gopher, Chinchilla also needs less compute for fine-tuning and inference. The author block prints no affiliations; the correspondence addresses are at deepmind.com, and the model card in the appendix names DeepMind as the developer.

Event record

Event date
March 29, 2022
Timeline date
Event date
Verification
Sources gathered automatically · September 24, 2026
Lines
ID
evt-0672

The only version of arXiv:2203.15556, 29 March 2022.

Sources

Related events

Records that link to this one