Back to timeline

Research · January 31, 2025

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

A thousand examples and a forced budget

On 31 January 2025 Stanford and the Allen Institute showed that reasoning switches on with a thousand examples. Fine-tuning Qwen2.5-32B-Instruct on the s1K set took 26 minutes on 16 H100 accelerators, and the model beat o1-preview on competition mathematics by up to 27 per cent.

Why it matters

Until then the ability to reason looked like a consequence of either enormous fine-tuning or long reinforcement learning. This work found the lower bound on its price: a thousand selected examples and half an hour on sixteen accelerators, which put reproducing such a model within reach of a university lab.

The s1K set is exactly one thousand problems with reasoning traces, selected on three criteria — difficulty, diversity and quality — each validated by ablation. The second device is budget forcing: if the model wants to stop early, the word "Wait" is appended and it keeps reasoning; if it reasons too long, the thinking is cut off. This is what lifts AIME24 from 50.0 to 56.7 per cent, which the abstract rounds to 57. For comparison, o1-preview scores 44.6 on AIME24. What the record does not claim. The authors name two limits on their own device: the gain from longer thinking eventually flattens out, and beyond that it runs into the base model's context window. Budget forcing stretches what the model already does rather than adding anything. The 27 per cent margin is on competition mathematics, MATH and AIME24, not on tasks in general.

Event record

Event date
January 31, 2025
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 22, 2026
Lines
ID
evt-0566

The day the first version of the preprint was submitted. arXiv later added two more; the figures were read in the first.

Sources

Related events

Records that link to this one