OpenAI assembles a superalignment team
On 5 July 2023 OpenAI announced a superalignment team led by Ilya Sutskever and Jan Leike and promised it 20% of the compute secured to date, over four years. The goal was to solve the core technical challenges of superintelligence alignment in four years by building a roughly human-level automated alignment researcher.
Why it matters
A laboratory publicly named the share of its resources it gave to safety and a deadline, which is a yardstick it could be checked against. Ten months later both leaders left, and one wrote that the team had been short of compute: that yardstick became the point of dispute.
The post: superintelligence could arrive this decade and could lead to the disempowerment of humanity or even extinction; current techniques such as reinforcement learning from human feedback rely on humans' ability to supervise and will not scale to superintelligence. The plan: scalable oversight with AI assistance, control of generalisation, automated search for problematic behaviour and internals, and deliberately training misaligned models to check that they are caught. The 20% is of the compute secured 'to date', that is as of July 2023, not of future compute. No document read establishes whether the promised compute was or was not provided.