A thousand examples and a forced budget
On 31 January 2025 Stanford and the Allen Institute showed that reasoning switches on with a thousand examples. Fine-tuning Qwen2.5-32B-Instruct on the s1K set took 26 minutes on 16 H100 accelerators, and the model beat o1-preview on competition mathematics by up to 27 per cent.
Why it matters
Until then the ability to reason looked like a consequence of either enormous fine-tuning or long reinforcement learning. This work found the lower bound on its price: a thousand selected examples and half an hour on sixteen accelerators, which put reproducing such a model within reach of a university lab.
The s1K set is exactly one thousand problems with reasoning traces, selected on three criteria — difficulty, diversity and quality — each validated by ablation. The second device is budget forcing: if the model wants to stop early, the word "Wait" is appended and it keeps reasoning; if it reasons too long, the thinking is cut off. This is what lifts AIME24 from 50.0 to 56.7 per cent, which the abstract rounds to 57. For comparison, o1-preview scores 44.6 on AIME24. What the record does not claim. The authors name two limits on their own device: the gain from longer thinking eventually flattens out, and beyond that it runs into the base model's context window. Budget forcing stretches what the model already does rather than adding anything. The 27 per cent margin is on competition mathematics, MATH and AIME24, not on tasks in general.