Coverage grows with the number of attempts
On 31 July 2024 a group from Stanford, Oxford and Google DeepMind showed that the share of problems solved grows with the number of generated attempts across four orders of magnitude. On SWE-bench Lite, DeepSeek-V2-Coder-Instruct rose from 15.9% at one sample to 56% at 250.
Why it matters
Until then compute at answering time meant a longer chain of reasoning inside a single attempt. The paper measured the other, simpler route — just repeat — and showed that 56% from a cheap model at 250 samples beats the 43% single-attempt state of the art on the costlier models of the day.
Coverage, the fraction of problems solved by any attempt, follows an exponentiated power law against the number of samples and holds over four orders of magnitude. On GSM8K and MATH, coverage with Llama-3 models passes 95% at 10,000 samples. The record does not claim that repetition solves a problem by itself. It converts into performance only where answers can be checked automatically — code and formal proofs. The authors state the limit plainly: without an automatic verifier, the usual ways of picking the right attempt out of many — majority voting and reward models — plateau beyond several hundred samples and fail to scale with the sample budget. The first author and Ronald Clark are at Oxford, the rest at Stanford and Google DeepMind; attributing the work to Stanford alone is inaccurate.