Proximal policy optimisation (PPO)
On 20 July 2017 John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford and Oleg Klimov of OpenAI posted proximal policy optimisation: a surrogate objective that clips the ratio of new to old policy probabilities (epsilon = 0.2 in the experiments), so that the same batch of samples can be used for several epochs of minibatch updates with first-order methods alone. On seven MuJoCo tasks of one million timesteps PPO beat the previous policy gradient methods almost everywhere; on 49 Atari games it won 30 by average reward over training, against 18 for ACER and one for A2C.
Why it matters
The stability of trust region policy optimisation became available without its second-order machinery, in a method compatible with shared policy and value networks and with dropout. InstructGPT in 2022 trained its models with the PPO algorithm (Schulman et al., 2017) and named them PPO and PPO-ptx; GRPO in 2024 describes itself as a variant of PPO.
PPO descends from trust region policy optimisation (TRPO; Schulman, Levine, Moritz, Jordan and Abbeel, University of California, Berkeley, arXiv:1502.05477, first version 19 February 2015), which the PPO paper calls relatively complicated and incompatible with architectures that include noise such as dropout or parameter sharing. The Atari result depends on the metric: by average reward over the last 100 episodes ACER wins 28 games and PPO 19. Figures are from the first version.