Back to timeline

Research · July 20, 2017

Proximal policy optimisation (PPO)

On 20 July 2017 John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford and Oleg Klimov of OpenAI posted proximal policy optimisation: a surrogate objective that clips the ratio of new to old policy probabilities (epsilon = 0.2 in the experiments), so that the same batch of samples can be used for several epochs of minibatch updates with first-order methods alone. On seven MuJoCo tasks of one million timesteps PPO beat the previous policy gradient methods almost everywhere; on 49 Atari games it won 30 by average reward over training, against 18 for ACER and one for A2C.

Why it matters

The stability of trust region policy optimisation became available without its second-order machinery, in a method compatible with shared policy and value networks and with dropout. InstructGPT in 2022 trained its models with the PPO algorithm (Schulman et al., 2017) and named them PPO and PPO-ptx; GRPO in 2024 describes itself as a variant of PPO.

PPO descends from trust region policy optimisation (TRPO; Schulman, Levine, Moritz, Jordan and Abbeel, University of California, Berkeley, arXiv:1502.05477, first version 19 February 2015), which the PPO paper calls relatively complicated and incompatible with architectures that include noise such as dropout or parameter sharing. The Atari result depends on the metric: by average reward over the last 100 episodes ACER wins 28 games and PPO 19. Figures are from the first version.

Event record

Event date
July 20, 2017
Timeline date
Event date
Verification
Sources gathered automatically · September 24, 2026
Lines
ID
evt-0667

The first of two versions of arXiv:1707.06347, 20 July 2017; the second is of 28 August 2017. The predecessor, TRPO (arXiv:1502.05477), is of 19 February 2015.

Sources

Related events