Back to timeline

Research · January 2, 2024

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

Self-play against one's own earlier state

On 2 January 2024 UCLA described SPIN: a model learns to tell its own answers from the previous iteration apart from the human answers in the same SFT set. The average score of zephyr-7b-sft-full on the Open LLM Leaderboard rose from 58.14 to 63.16 over three iterations, with no new annotation.

Why it matters

Until then getting past the plateau after SFT meant either new human annotation or answers from a stronger commercial model as teacher. SPIN showed a third source of signal — contrast with one's own earlier state — and it turned out cheaper than both.

The iterations from the first version's table: 58.14, 60.80, 62.26, 62.97, 63.16, so the gain falls off quickly. On MT-Bench, 5.94 to 6.78. The largest movers were GSM8k (26.76 to 38.97) and TruthfulQA (43.73 to 54.90), over ten points each. Only a 50-thousand subset of Ultrachat200k was used. Against DPO: iteration zero of self-play is comparable to DPO trained on an additional 62 thousand preference pairs, and from iteration one self-play surpasses DPO on the majority of the sets. What the record does not claim. This is not training without human data: the reference the model learns to distinguish itself from is the human answers in the existing SFT set. No new annotation is needed; the existing annotation is.

Event record

Event date
January 2, 2024
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0534

The day the first version of the preprint was submitted. Later versions followed in February and June 2024; the figures were read in the first.

Sources

Related events