Back to timeline

Benchmark · April 14, 2026

A human baseline is measured for ARC-AGI-3

ARC Prize ran a study with 458 participants and changed the normalisation: 100 per cent now means clearing every level at or above median human action efficiency.

Why it matters

The denominator without which later percentages on this benchmark meant nothing now exists, measured on people under the same first-run conditions the models face.

The study was published on 14 April 2026: 458 participants, weekly in-person sessions of 90 minutes at a San Francisco testing centre, under first-run conditions in which each participant sees an environment only once. Base payment was about $130 plus $5 for each environment solved. Every environment was beaten by at least two independent participants and most by many more. Across the 25 public demo environments the solve rate ranges from 100 per cent, 10 solves in 10 plays on task r11l, down to 14 per cent, 2 in 14 on bp35. The normalising baseline moved from the second-best player to the median player per level: from now on a score of 100 per cent for AI means beating every level of every environment at or above median human action efficiency. Human and machine participants received identical information.

Event record

Event date
April 14, 2026
Timeline date
Event date
Verification
Sources gathered automatically · September 19, 2026
Lines
ID
evt-0379

The day ARC Prize published the study.

Sources

Related events

Records that link to this one