Five concrete problems in AI safety
On 21 June 2016 Dario Amodei and Chris Olah (Google Brain), Jacob Steinhardt (Stanford), Paul Christiano (Berkeley), John Schulman (OpenAI) and Dan Mané (Google Brain) posted the preprint 'Concrete Problems in AI Safety'. They reduced the risk of accidents in machine learning systems to five research problems: side effects, reward hacking, scalable oversight, safe exploration and distributional shift.
Why it matters
AI safety was restated in the language machine learning works in: not a future superintelligence but today's agents with a wrongly specified objective, expensive supervision or unfamiliar data. The very next year the paper from which learning from human preferences grew cites the list on misaligned objectives and on an agent that maximises only part of the true reward.
The paper sorts the problems by origin: a wrong objective (avoiding side effects and reward hacking), an objective too expensive to evaluate often (called 'scalable supervision' in the first version's abstract and 'scalable oversight' in the body), and undesirable behaviour during learning (safe exploration, distributional shift). The running example is an office cleaning robot; among the reward-hacking examples is a robot rewarded for the bleach it uses, which pours it down the drain. The work is a survey and a research agenda, not an experiment: it gives no figures about systems. Affiliations and date are from the first version.