Back to timeline

Benchmark · January 15, 2026

The ARC Prize 2025 technical report

The report recorded the competition: 1,455 teams, 15,154 entries and a top score of 24 per cent on the ARC-AGI-2 private set. It named the refinement loop as the theme of the year and the benchmark's own adoption as a source of new contamination.

Why it matters

The year closed with a number to measure from, and with the authors' claim that the yardstick is being degraded by the very adoption that made it standard.

The authors are Francois Chollet, Mike Knoop, Gregory Kamradt and Bryan Landers. The Kaggle competition attracted 1,455 teams and 15,154 entries, with a top score of 24 per cent on the ARC-AGI-2 private evaluation set. Paper submissions nearly doubled year over year to 90 entries. The defining theme of 2025 is named as the emergence of the refinement loop: a per-task iterative program optimisation loop guided by a feedback signal. Four frontier labs - Anthropic, Google DeepMind, OpenAI and xAI - reported ARC-AGI performance in public model cards during 2025, which made the benchmark an industry standard. The authors argue at the same time that frontier reasoning performance remains fundamentally constrained by knowledge coverage, and that this gives rise to new forms of benchmark contamination.

Event record

Event date
January 15, 2026
Timeline date
Event date
Verification
Sources gathered automatically · September 19, 2026
Lines
ID
evt-0366

The day the report was submitted to arXiv.

Sources

Related events

Records that link to this one