Back to timeline

Research · May 23, 2025

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

A small model is handed actions, not answers

On 23 May 2025 KAIST and KRAFTON showed agent distillation: a small model takes from a larger one not reasoning chains but trajectories of action, with retrieval and code execution. Models of 0.5, 1.5 and 3 billion parameters matched models one tier up — 1.5, 3 and 7 billion.

Why it matters

Until then transferring skill from a large model to a small one meant copying what the teacher wrote. Here what is copied is what the teacher did: when and where it reached for information and what code it ran. The small model ends up with not the teacher's knowledge but its way of obtaining knowledge.

Two devices. The first is the first-thought prefix: the teacher is nudged to begin with a thought, which improves the quality of the trajectories it generates. The second is self-consistent action generation, which adds robustness at answering time, when the small model acts on its own. The comparison is not against a base model but against the same size trained by ordinary chain-of-thought distillation: 0.5 billion with agent distillation runs level with 1.5 billion on chains, and likewise at the two other tiers. Tested on eight reasoning tasks, factual and mathematical, covering both in-domain and out-of-domain problems. What the record does not claim. The gain is described as being competitive with the next tier up, not as beating it. The record is not dated 20 May as the report had it: arXiv's submission history gives 23 May, and no version falls on the 20th.

Event record

Event date
May 23, 2025
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 22, 2026
Lines
ID
evt-0576

The day the first version was submitted, by arXiv's submission history. The report gave 20 May; no version carries that date.

Sources

Related events