Scaling laws
The study measured loss relative to scale.
Why it matters
It made scale a measurable design variable for language-model training.
The relationships reported are empirical fits over the studied training range.
Research · January 23, 2020
The study measured loss relative to scale.
It made scale a measurable design variable for language-model training.
The relationships reported are empirical fits over the studied training range.
arXiv · Published January 23, 2020
The study measured models of the same architectural family.
Both records concern computing resources. Training scaling laws do not establish the lineage of this inference chip.
Computation on that scale made the scaling experiments possible.
The model size was chosen from the scaling laws published the same year.
Language Models are Few-Shot LearnersQuality grows with compute spent at answering time, not only at training time.
The scaling laws described quality rising smoothly with compute; this set looked for tasks where it does not, and named brittle metrics as one cause of the jumps.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language modelsThe scaling law says how much compute is worth spending; this series says how much was actually spent and how fast that grew.
Compute Trends Across Three Eras of Machine Learning (arXiv:2202.05924v1)Kaplan and co-authors advised, for a tenfold budget, a 5.5-fold larger model on only 1.8 times more tokens; this paper finds that model size and training tokens should be scaled in equal proportions.
Training Compute-Optimal Large Language Models