Gold at the International Mathematical Olympiad
Two laboratories independently scored 35 points out of 42, solving five problems of six in ordinary language, with no formal systems and no tools.
Why it matters
A year earlier silver needed a specialised system with machine checking; now gold went to a general-purpose model simply reasoning in text.
AlphaProof in 2024 translated the problem into Lean, where a machine checks every step. Here the models wrote proofs in natural language, as a person does, and a human jury graded them under the same rules. They had the same time as the contestants, four and a half hours. Neither managed the sixth problem. The difference in approach matters: formal checking gave a guarantee of correctness and ordinary language does not, so reliability now rests on the model itself.