Compute at answering time against a bigger model
On 6 August 2024 four researchers from Berkeley and Google DeepMind measured what compute at answering time is worth. On problems where a smaller model already has some success, that model with extra answering-time compute beat a model 14 times larger at the same total operation count.
Why it matters
Until then the link between more thinking and better answers was known from practice but had not been measured against the alternative of spending the same operations on model size. The paper put both on one FLOPs axis and showed where the limit is: where the base model has no success at all, training still wins.
Two mechanisms were measured: search against a process reward model (PRM), and sequentially revising the model's own draft. A strategy that picks the mechanism from the difficulty of the question gained a factor of 2-4 in efficiency over plain best-of-N — the paper says "2-4x" and "up to 4x less test-time compute", not more. The dependence on difficulty runs the opposite way to the obvious guess: on easier problems, where the model already produces a reasonable answer, sequential revision works better, while on harder ones independent parallel sampling or tree search against a PRM does. The record does not claim that answering-time compute replaces training. The authors state the opposite for the hardest questions: there the extra operations are better spent on pretraining, so the exchange is not one for one. The comparison against the 14-times-larger model holds under the stated condition that the smaller model already attains "somewhat non-trivial success rates" on those problems; outside that condition it is not claimed.