Back to timeline

Research · February 4, 2016

Asynchronous actor-critic (A3C)

On 4 February 2016 Volodymyr Mnih and co-authors at Google DeepMind posted asynchronous variants of four reinforcement learning algorithms, in which many actor-learners run in parallel on copies of the environment instead of learning from a replay memory. The best, asynchronous advantage actor-critic (A3C), after four days on 16 CPU cores with no GPU reached a mean of 623.0 and a median of 112.6 per cent of human-normalised score on 57 Atari games, against 121.9 and 47.5 per cent for DQN after eight days on a GPU.

Why it matters

Deep reinforcement learning no longer needed a replay memory to be stable: parallel actors decorrelate the data by themselves, so on-policy methods such as actor-critic became usable with deep networks, and training moved from a GPU to a multi-core CPU. The synchronous form, A2C, is what the 2017 work on learning from human preferences used for Atari and what PPO compared itself with.

Table 1, read from the rendered page: feed-forward A3C after one day on CPU 344.1 per cent mean and 68.2 median, after four days 496.8 and 116.6; A3C with an LSTM after four days 623.0 and 112.6; Gorila, four days on 100 machines, 215.2 and 71.3; Prioritized DQN, eight days on GPU, 463.6 and 127.6. The comparison agents were trained for 8 to 10 days on Nvidia K40 GPUs. The paper also shows TORCS car racing, MuJoCo continuous control and a new task of finding rewards in random 3D mazes (Labyrinth) from pixels. One of the eight authors, Mehdi Mirza, was at MILA, University of Montreal. What the record does not claim: that A3C beat every agent on the median, where Prioritized DQN and Dueling Double DQN (117.1 per cent) are higher.

Event record

Event date
February 4, 2016
Timeline date
Event date
Verification
Sources gathered automatically · September 24, 2026
Lines
ID
evt-0664

The first of two versions of arXiv:1602.01783, 4 February 2016; the second is of 16 June 2016. Figures are from the first.

Sources

Related events

Records that link to this one