Back to timeline

Benchmark · March 5, 2021

MATH: competition mathematics

On 5 March 2021 Dan Hendrycks and colleagues at Berkeley and the University of Chicago released MATH, 12,500 problems from school mathematics competitions (7,500 training and 5,000 test) with step-by-step solutions, in seven subjects and five levels of difficulty. Large language models solved between 2.9 and 6.9%.

Why it matters

The set showed that scale alone does not close mathematics: accuracy rose slowly with model size, and the authors estimated that on a log-linear trend 40% would take around 10^35 parameters. MATH later became one of the sets on which models' reasoning is measured, including in records of the atlas from 2024 and 2025.

With the set the authors assembled AMPS, an auxiliary pretraining corpus: Khan Academy problems and about 5 million problems generated by 100 Mathematica scripts, over 23 GB. People were tested too, on 20 problems: a computer science PhD student who does not like mathematics scored about 40%, a three-time International Mathematical Olympiad gold medallist 90%. The 6.9% came from a 1.5-billion-parameter GPT-2 pretrained on AMPS and fine-tuned on MATH.

Event record

Event date
March 5, 2021
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0789

The first version of the preprint.

Sources

Related events