Five thousand puzzles instead of a corpus
On 20 February 2025 researchers at Microsoft Research Asia showed that rule-based reinforcement learning works on a very small synthetic corpus. After five thousand generated logic puzzles, a 7-billion-parameter model improved by 125 per cent on AIME and 38 per cent on AMC.
Why it matters
DeepSeek-R1 had shown that reasoning emerges from reinforcement, but on closed data and at scale. Here the same is reproduced on a corpus that can be generated procedurally and checked by rule — and the ability turned out to carry over to mathematics, which the corpus did not contain at all.
The corpus is procedurally generated knights-and-knaves puzzles, where the answer is checked by a rule rather than by a person or a model. That is what makes the reward here genuinely rule-based. The interesting part is what appeared outside the corpus. The model acquired reflection, self-verification and summarisation, none of which is present in the logic corpus. The transfer to AIME and AMC, that is to mathematics competitions, is what the paper reads as evidence that reinforcement develops abstract problem-solving schemata rather than memorising a domain. What the record does not claim. The 125 and 38 per cent are gains relative to the base model, not absolute accuracy and not a comparison against frontier systems. One scale was tested: 7 billion parameters.