Back to timeline

Research · May 29, 2023

Direct preference optimisation

The authors showed that alignment to human preferences reduces to an ordinary loss function, with no reward model and no reinforcement learning.

Why it matters

The hardest part of the ChatGPT recipe simplified to a few lines of code, and alignment became available to small teams.

The InstructGPT scheme required three stages and the tuning of reinforcement learning, which is known for instability. DPO proves that the optimal policy can be expressed through the reward model analytically, so the reward need not be built at all. The method quickly became standard for open models because it is reproducible and cheap.

Event record

Event date
May 29, 2023
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0277

Preprint of 29 May 2023; presented at NeurIPS 2023.

Sources

Related events