Direct preference optimisation
The authors showed that alignment to human preferences reduces to an ordinary loss function, with no reward model and no reinforcement learning.
Why it matters
The hardest part of the ChatGPT recipe simplified to a few lines of code, and alignment became available to small teams.
The InstructGPT scheme required three stages and the tuning of reinforcement learning, which is known for instability. DPO proves that the optimal policy can be expressed through the reward model analytically, so the reward need not be built at all. The method quickly became standard for open models because it is reproducible and cheap.