Back to timeline

Research · June 12, 2017

Learning from human preferences

On 12 June 2017 Paul Christiano and Dario Amodei of OpenAI, Jan Leike, Miljan Martic and Shane Legg of DeepMind, and Tom Brown posted a method in which a person compares pairs of short segments of an agent's trajectory, a reward model is fitted to those choices, and the agent learns by reinforcement on the predicted reward. On eight MuJoCo robotics tasks and seven Atari games it worked without access to the reward function, with human feedback on less than 1 per cent of the agent's interactions with the environment.

Why it matters

A goal that cannot be written as a reward function could be communicated through comparisons made by non-experts, at a cost the authors put at about an hour of human time for a new behaviour. This is the reinforcement learning from human feedback that InstructGPT cites (Christiano et al., 2017) when it fine-tunes GPT-3 on human preferences in 2022.

With 700 queries to a human rater the method nearly matched reinforcement learning on the true reward in the MuJoCo tasks; the Atari games used 5,500 queries. The Hopper robot learned a sequence of backflips, a behaviour with no reward function, from 900 queries in less than an hour. Policies were trained with TRPO for robotics and with A2C, the synchronous form of A3C, for Atari. Tom Brown is listed in the author block without an institution.

Event record

Event date
June 12, 2017
Timeline date
Event date
Verification
Sources gathered automatically · September 24, 2026
Lines
ID
evt-0666

The first of four versions of arXiv:1706.03741, 12 June 2017; the last is of 17 February 2023. Figures are from the first.

Sources

Related events

Records that link to this one