Back to timeline

Research · March 2022

InstructGPT and learning from human ratings

OpenAI fine-tuned a model to carry out requests by training a separate reward model on human comparisons of answers.

Why it matters

A 1.3-billion-parameter model became more useful than a 175-billion one: alignment mattered more than size.

The scheme has three steps: fine-tuning on examples of desired behaviour, training a reward model on human rankings of answers, then optimising the policy against that reward. The result was not about the quality of the language but about whether the model does what it is asked. Eight months later the same scheme turned GPT-3.5 into ChatGPT, and it is this, not size, that made the model broadly usable.

Event record

Event date
March 2022
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0243

Preprint of 4 March 2022; presented at NeurIPS 2022.

Sources

Related events

Records that link to this one