Back to timeline

Benchmark · October 27, 2021

GSM8K: school problems and a verifier

On 27 October 2021 Karl Cobbe and colleagues at OpenAI released GSM8K, 8.5 thousand grade-school maths word problems (7.5 thousand for training and a thousand for testing), each of 2 to 8 steps. With the set they proposed a verifier that scores the model's candidate solutions and picks the best.

Why it matters

The problems are conceptually simple, yet even the largest models did not solve them reliably, so the set measured multi-step reasoning as such. The verifier gave about the same gain as a 30-fold increase in model size, an early measurement that computation at answer time can partly stand in for size.

The authors extrapolated that 80% would need a model of around 10^16 parameters. The first thousand problems were written by freelancers hired through Upwork, the rest through the labelling platform Surge AI; each problem was then re-solved by another worker, disagreements were repaired or discarded, and a sample check left 1.7% of problems disputed. The record gives no accuracy for individual models: in the preprint they appear only in figures.

Event record

Event date
October 27, 2021
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0790

The first version of the preprint.

Sources

Related events