Back to timeline

Research · November 19, 2019

MuZero: planning without the rules

On 19 November 2019 Julian Schrittwieser and colleagues at DeepMind described MuZero: tree search over a learned model that is not given the rules of the game. On 57 Atari games it set a new best result, and at Go, chess and shogi it matched AlphaZero, which knew the rules.

Why it matters

Planning by tree search no longer needed an exact simulator of the environment. The model learns to predict only what planning needs, reward, policy and value, and so the same method works in board games and in visually complex video games.

By the first preprint version MuZero beat the previous best method, R2D2, in 42 of 57 Atari games on mean and median normalised score, and the model-based SimPLe in all of them; at Go it slightly exceeded AlphaZero. The Nature paper was published on 23 December 2020; its text was not read (it is behind a subscription), so the record claims only what the preprint says.

Event record

Event date
November 19, 2019
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0787

The first version of the preprint; the Nature paper appeared on 23 December 2020.

Sources

Related events