GSM8K: school problems and a verifier
On 27 October 2021 Karl Cobbe and colleagues at OpenAI released GSM8K, 8.5 thousand grade-school maths word problems (7.5 thousand for training and a thousand for testing), each of 2 to 8 steps. With the set they proposed a verifier that scores the model's candidate solutions and picks the best.
Why it matters
The problems are conceptually simple, yet even the largest models did not solve them reliably, so the set measured multi-step reasoning as such. The verifier gave about the same gain as a 30-fold increase in model size, an early measurement that computation at answer time can partly stand in for size.
The authors extrapolated that 80% would need a model of around 10^16 parameters. The first thousand problems were written by freelancers hired through Upwork, the rest through the labelling platform Surge AI; each problem was then re-solved by another worker, disagreements were repaired or discarded, and a sample check left 1.7% of problems disputed. The record gives no accuracy for individual models: in the preprint they appear only in figures.