InstructGPT and learning from human ratings
OpenAI fine-tuned a model to carry out requests by training a separate reward model on human comparisons of answers.
Why it matters
A 1.3-billion-parameter model became more useful than a 175-billion one: alignment mattered more than size.
The scheme has three steps: fine-tuning on examples of desired behaviour, training a reward model on human rankings of answers, then optimising the policy against that reward. The result was not about the quality of the language but about whether the model does what it is asked. Eight months later the same scheme turned GPT-3.5 into ChatGPT, and it is this, not size, that made the model broadly usable.