DeepSeek-R1
A Chinese laboratory opened a model that reasons at the level of o1 and showed that the behaviour emerges from pure reinforcement learning on the correct answer.
Why it matters
Reasoning stopped being one company's closed technology: the recipe turned out to be short and the weights were in the open.
What mattered was not the result but the method. The base model was trained by reinforcement on problems with a checkable answer, with no human example of reasoning at all, and it developed long chains with self-checking and backtracking by itself. The report described the moment during training when the model stops and rewrites its approach. Unlike o1, the chain was shown in full. The weights were published under an MIT licence, and dozens of smaller models grew on them within weeks.