Back to timeline

Research · December 15, 2022

Constitutional AI: harmlessness from AI feedback

On 15 December 2022, 51 authors from Anthropic led by Yuntao Bai posted the preprint 'Constitutional AI'. An assistant was trained to be harmless without a single human label identifying harmful outputs: human oversight was reduced to a short list of principles, a 'constitution', and the comparisons from which the preference model learns were made by the model itself.

Why it matters

Learning from human feedback typically used at least tens of thousands of human labels; here the harmlessness labels were replaced by the model's own judgements, and the method was named RLAIF, reinforcement learning from AI feedback. The result is an assistant that does not evade a harmful request but explains its objections. A year later preference models built with the method were used in the safety training of Sleeper Agents.

Two phases: in supervised learning the model critiques and revises its own answers and is fine-tuned on the revisions; in the reinforcement phase the model picks the better of two answers, a preference model is trained on these comparisons, and it becomes the reward signal. Both phases can use chain-of-thought reasoning. There are 'of order ten' principles, which the authors say were chosen in a fairly ad hoc and iterative way for research purposes and ought in future to be redeveloped by a wider set of stakeholders. Helpfulness was still trained on human judgements; only harmlessness went without human labels. The figure of 16 principles that turns up in retellings is not in the paper.

Event record

Event date
December 15, 2022
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0717

The day of the only version of arXiv:2212.08073.

Sources

Related events

Records that link to this one