Self-play against one's own earlier state
On 2 January 2024 UCLA described SPIN: a model learns to tell its own answers from the previous iteration apart from the human answers in the same SFT set. The average score of zephyr-7b-sft-full on the Open LLM Leaderboard rose from 58.14 to 63.16 over three iterations, with no new annotation.
Why it matters
Until then getting past the plateau after SFT meant either new human annotation or answers from a stronger commercial model as teacher. SPIN showed a third source of signal — contrast with one's own earlier state — and it turned out cheaper than both.
The iterations from the first version's table: 58.14, 60.80, 62.26, 62.97, 63.16, so the gain falls off quickly. On MT-Bench, 5.94 to 6.78. The largest movers were GSM8k (26.76 to 38.97) and TruthfulQA (43.73 to 54.90), over ten points each. Only a 50-thousand subset of Ultrachat200k was used. Against DPO: iteration zero of self-play is comparable to DPO trained on an additional 62 thousand preference pairs, and from iteration one self-play surpasses DPO on the majority of the sets. What the record does not claim. This is not training without human data: the reference the model learns to distinguish itself from is the human answers in the existing SFT set. No new annotation is needed; the existing annotation is.