Chain-of-thought prompting
On 28 January 2022 Jason Wei and six colleagues at Google Brain posted chain-of-thought prompting: eight worked examples in the prompt show the intermediate steps before each answer, and the model then writes such steps itself. With a 137-billion-parameter model, accuracy on the GSM8K grade-school maths problems rose from 6.3 to 14.8 per cent, and to 19.5 with an external calculator, against 18 per cent for GPT-3 175B fine-tuned with a calculator; on MultiArith, from 7.6 to 45.0.
Why it matters
Reasoning could be drawn out of a large model by the form of the prompt, without training, and the gain appeared only at about 100 billion parameters: smaller models wrote fluent but illogical chains and did worse than with standard prompts. In 2024 o1 makes the chain of thought the thing reinforcement learning trains.
The first version tests a 137-billion-parameter model (LaMDA-PT, trained on 2.49 trillion tokens) and GPT-3 davinci. PaLM 540B and the 58 per cent on GSM8K that are usually quoted for this paper do not occur in the first version at all and came with later versions. Of 50 random chains that led to a wrong answer, 46 per cent were almost correct and 54 per cent had major errors of understanding or coherence. Ablations for the 137B model show that the gain comes neither from writing the equation alone, nor from spending more tokens, nor from reasoning written after the answer.