GRPO: reinforcement learning without a critic
On 5 February 2024 DeepSeek described Group Relative Policy Optimization, a variant of PPO that drops the critic model and takes its baseline from the spread of scores across a group of answers to the same question. Accuracy on MATH rose from 46.8% to 51.7%.
Why it matters
The critic in PPO is a second network the size of the model itself, and it was what made reinforcement learning expensive. Removing it made cheaper the kind of post-training with which DeepSeek-R1 would be trained a year later.
Figures from the first version of the preprint: the base model DeepSeekMath-Base 7B scores 36.2% on MATH and with that beats the closed Minerva 540B, a model 77 times larger. After SFT (DeepSeekMath-Instruct 7B) it is 46.8%, after the reinforcement phase under GRPO (DeepSeekMath-RL 7B) 51.7%, and on GSM8K 82.9% to 88.2%. The paper says this "approaches" the level of GPT-4 and Gemini-Ultra rather than matching it. What the record does not claim. This work did not introduce a reward a program checks: the word "verifiable" does not occur once in the text, and the reinforcement phase uses a reward model trained on DeepSeekMath-Base 7B. GRPO removes the critic — the value estimator inside PPO — not the reward model; these are different components, and confusing them credits the paper with someone else's result. A rule in place of a reward model comes later, with DeepSeek-R1. The paper names no figure for the memory saved: it says only that the critic is "a model of comparable size as the policy model" and that dropping it notably reduces the cost.