The cost of an answer is memory, not parameters
On 5 June 2025 Carnegie Mellon recalculated the scaling laws for answering-time compute with hardware taken into account and reached the opposite conclusion: the usefulness of small models is significantly overestimated, and resources should first go into model size, up to about 14 billion parameters.
Why it matters
Earlier work that year counted computation and concluded that a small model thinking for longer beats a large one. This recalculation added what those counts left out — memory access in the attention mechanism — and that, not parameter count, turned out to govern cost.
The analysis spans models from 0.6 to 32 billion parameters and strategies such as best-of-N and long chains of thought. The conclusion is put as an order of spending: grow the model to a critical threshold first, empirically around 14 billion, and only then invest in answering-time compute. A second consequence follows: if attention is the bottleneck, sparse attention lowers per-token cost and allows either longer generations or more parallel samples within the same budget. Sparse models give over 60 points of accuracy on AIME in the low-cost regime and over 5 points in the high-cost one; by the paper's accounting that is up to three times fewer resources for the same accuracy. What the record does not claim. The 14-billion threshold is called empirical rather than derived, and it applies to the models and strategies tested. The over-60-point gain belongs to the low-cost regime on AIME and is not a typical figure: in the high-cost regime it falls to over 5.