Back to timeline

Research · February 17, 2025

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

Without a verifier, answering-time compute loses

On 17 February 2025 Carnegie Mellon and Berkeley proved that at equal budget, fine-tuning with a verifier beats imitating someone else's reasoning traces, and that the gap between them widens as answering-time compute grows.

Why it matters

At that moment the cheapest way to build a reasoning model was to copy finished traces from a stronger one. The work showed that cheapness has a ceiling: without a verification signal the gain from longer thinking runs out sooner, and the more compute is spent the worse the deficit.

Two classes are separated. Verifier-based methods collect reward annotations for the base model's own rollouts; verifier-free methods carry over someone else's reasoning traces without querying any verification signal. The paper places s1 and OpenThinker in that second class by name. The condition under which the advantage holds is stated openly: the base model must be heterogeneous in its answers, meaning the distribution of rewards over its own traces must not be sharp. The authors formalise this as anti-concentration. Under it, verifier-based methods gain a factor of the square root of H in sample efficiency, where H is the output length. Verified both theoretically and on models of 3, 8 and 32 billion parameters, on a didactic task and on real mathematical ones. What the record does not claim. The advantage is not unconditional: it is proved under anti-concentration, and the paper stresses this rather than presenting it as a general rule. The work does not measure shipped products, and it does not say distillation is useless — it says distillation scales worse.

Event record

Event date
February 17, 2025
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 22, 2026
Lines
ID
evt-0568

The day the first version of the preprint was submitted. arXiv later added a second; the figures were read in the first.

Sources

Related events

Records that link to this one