One billion against four hundred and five
On 10 February 2025 Shanghai AI Laboratory and Tsinghua measured how to allocate answering-time compute properly. A one-billion-parameter model beat a 405-billion one on MATH-500, and a seven-billion model beat o1 and DeepSeek-R1.
Why it matters
Until then answering-time compute looked like a way to pull a model up a little. Here it turned out that allocated well it covers a two- or three-order difference in size — so the question of which model is needed depends on how long it is allowed to think.
Measured on MATH-500 and AIME24, with reward models from 1.5 to 72 billion parameters and several base models. Four comparisons from the first version: a 1-billion model beats a 405-billion one on MATH-500; a 0.5-billion model beats GPT-4o on both benchmarks; a 3-billion model beats the 405-billion one on both; a 7-billion model beats o1 and DeepSeek-R1, and is cheaper at inference. What the record does not claim. The first of the paper's own conclusions, and the one the report about it leaves out, is that the optimal allocation depends heavily on three things at once: the base model, the reward model and the difficulty of the problem. This is not a universal device to be switched on anywhere but a result of matching a particular trio. The comparisons are on mathematical benchmarks, not on tasks in general.