Back to timeline

Research · January 22, 2025

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

Long context inside reinforcement learning

On 22 January 2025 Moonshot AI described Kimi k1.5, in which the context window is stretched to 128 thousand tokens not at inference but inside the reinforcement learning phase itself. The model scores 77.5 on AIME and 96.2 on MATH-500, matching o1.

Why it matters

Until then long context was a property of a finished model. Here it became a condition of training: the longer the model can reason during the reinforcement itself, the better it learns, and the report shows quality rising with that length.

Long-chain figures from the first version: 77.5 on AIME, 96.2 on MATH-500, 94th percentile on Codeforces, 74.9 on MathVista. The model reasons jointly over text and images, so MathVista is not a stray number here. A separate result is long2short: long-chain techniques are carried into short-output models through a length penalty, model merging, shortest rejection sampling and DPO. That gives 60.8 on AIME, 94.6 on MATH500 and 47.3 on LiveCodeBench, which the report calls a margin of up to 550 per cent over GPT-4o and Claude Sonnet 3.5. To make training at that window possible, partial rollouts were introduced: a long rollout is not restarted from nothing each time. Training runs in two epochs, first at 32 thousand tokens and then at 128 thousand. What the record does not claim. The report names no individual authors; it is signed "Kimi Team". The 550 per cent is an upper bound on one benchmark, not a typical margin.

Event record

Event date
January 22, 2025
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 22, 2026
Lines
ID
evt-0567

The day the first version of the technical report was submitted. arXiv later added three more; the figures were read in the first.

Sources

Related events