Constitutional AI: harmlessness from AI feedback
On 15 December 2022, 51 authors from Anthropic led by Yuntao Bai posted the preprint 'Constitutional AI'. An assistant was trained to be harmless without a single human label identifying harmful outputs: human oversight was reduced to a short list of principles, a 'constitution', and the comparisons from which the preference model learns were made by the model itself.
Why it matters
Learning from human feedback typically used at least tens of thousands of human labels; here the harmlessness labels were replaced by the model's own judgements, and the method was named RLAIF, reinforcement learning from AI feedback. The result is an assistant that does not evade a harmful request but explains its objections. A year later preference models built with the method were used in the safety training of Sleeper Agents.
Two phases: in supervised learning the model critiques and revises its own answers and is fine-tuned on the revisions; in the reinforcement phase the model picks the better of two answers, a preference model is trained on these comparisons, and it becomes the reward signal. Both phases can use chain-of-thought reasoning. There are 'of order ten' principles, which the authors say were chosen in a fairly ad hoc and iterative way for research purposes and ought in future to be redeveloped by a wider set of stakeholders. Helpfulness was still trained on human judgements; only harmlessness went without human labels. The figure of 16 principles that turns up in retellings is not in the paper.