Back to timeline

Research · February 5, 2024

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

GRPO: reinforcement learning without a critic

On 5 February 2024 DeepSeek described Group Relative Policy Optimization, a variant of PPO that drops the critic model and takes its baseline from the spread of scores across a group of answers to the same question. Accuracy on MATH rose from 46.8% to 51.7%.

Why it matters

The critic in PPO is a second network the size of the model itself, and it was what made reinforcement learning expensive. Removing it made cheaper the kind of post-training with which DeepSeek-R1 would be trained a year later.

Figures from the first version of the preprint: the base model DeepSeekMath-Base 7B scores 36.2% on MATH and with that beats the closed Minerva 540B, a model 77 times larger. After SFT (DeepSeekMath-Instruct 7B) it is 46.8%, after the reinforcement phase under GRPO (DeepSeekMath-RL 7B) 51.7%, and on GSM8K 82.9% to 88.2%. The paper says this "approaches" the level of GPT-4 and Gemini-Ultra rather than matching it. What the record does not claim. This work did not introduce a reward a program checks: the word "verifiable" does not occur once in the text, and the reinforcement phase uses a reward model trained on DeepSeekMath-Base 7B. GRPO removes the critic — the value estimator inside PPO — not the reward model; these are different components, and confusing them credits the paper with someone else's result. A rule in place of a reward model comes later, with DeepSeek-R1. The paper names no figure for the memory saved: it says only that the critic is "a model of comparable size as the policy model" and that dropping it notably reduces the cost.

Event record

Event date
February 5, 2024
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0528

The day the first version of the preprint was submitted. The second followed the next day and the third on 27 April 2024; the figures were read in the first.

Sources

Related events

Records that link to this one