Back to timeline

Benchmark · September 1, 2026

FrontierMath Erdos: 68 unsolved problems as a benchmark

Epoch AI assembled a benchmark of 68 open Erdos problems formalised in Lean. Of five models tested, only GPT-6 Astra solved anything, and it solved two.

Why it matters

The informal scoreboard of which Erdos problem fell this week was replaced by a fixed set with a budget, machine checking and a cost per solution.

The publication is dated 1 September 2026. The benchmark consists of 68 significant Erdos problems, open as of August 2026, selected by the mathematician Thomas Bloom and formalised in Lean. Each model received one attempt per problem with a budget of 300 dollars and 72 hours of working time. Only GPT-6 Astra, in a pre-release version, solved anything: two of the 68 problems, or 3 per cent. It disproved problem 74 by finding a counterexample, at a cost of 222 dollars, and proved problem 126 at a cost of 172 dollars. GPT-5.6 Sol, GPT-5.5, Claude Fable 5.1 and Claude Fable 5 each scored zero. Epoch notes that reaching these five solutions took over 220 thousand dollars of compute across all attempts.

Event record

Event date
September 1, 2026
Timeline date
Event date
Verification
Sources gathered automatically · September 19, 2026
Lines
ID
evt-0405

The day Epoch AI announced the benchmark.

Sources

Related events