Chinchilla: compute-optimal training
On 29 March 2022 Jordan Hoffmann and colleagues at DeepMind posted the result of training over 400 language models from 70 million to over 16 billion parameters on 5 to 500 billion tokens: for compute-optimal training, model size and the number of tokens should be scaled equally. Chinchilla, 70 billion parameters trained on 1.4 trillion tokens with Gopher's budget (4 times more data than the 280-billion-parameter Gopher), outperformed Gopher, GPT-3, Jurassic-1 and Megatron-Turing NLG and reached 67.5 per cent on MMLU, a greater than 7 per cent improvement over Gopher.
Why it matters
The rule by which large models were sized was rewritten: by Kaplan and co-authors (2020) a tenfold budget meant a 5.5-fold larger model on only 1.8 times more tokens, and most large models had been trained on about 300 billion tokens. By this paper they were significantly undertrained. LLaMA in 2023 takes this recommendation as its point of departure.
Every doubling of model size should come with a doubling of training tokens. Being smaller than Gopher, Chinchilla also needs less compute for fine-tuning and inference. The author block prints no affiliations; the correspondence addresses are at deepmind.com, and the model card in the appendix names DeepMind as the developer.