MuZero: planning without the rules
On 19 November 2019 Julian Schrittwieser and colleagues at DeepMind described MuZero: tree search over a learned model that is not given the rules of the game. On 57 Atari games it set a new best result, and at Go, chess and shogi it matched AlphaZero, which knew the rules.
Why it matters
Planning by tree search no longer needed an exact simulator of the environment. The model learns to predict only what planning needs, reward, policy and value, and so the same method works in board games and in visually complex video games.
By the first preprint version MuZero beat the previous best method, R2D2, in 42 of 57 Atari games on mean and median normalised score, and the model-based SimPLe in all of them; at Go it slightly exceeded AlphaZero. The Nature paper was published on 23 December 2020; its text was not read (it is behind a subscription), so the record claims only what the preprint says.