A human baseline is measured for ARC-AGI-3
ARC Prize ran a study with 458 participants and changed the normalisation: 100 per cent now means clearing every level at or above median human action efficiency.
Why it matters
The denominator without which later percentages on this benchmark meant nothing now exists, measured on people under the same first-run conditions the models face.
The study was published on 14 April 2026: 458 participants, weekly in-person sessions of 90 minutes at a San Francisco testing centre, under first-run conditions in which each participant sees an environment only once. Base payment was about $130 plus $5 for each environment solved. Every environment was beaten by at least two independent participants and most by many more. Across the 25 public demo environments the solve rate ranges from 100 per cent, 10 solves in 10 plays on task r11l, down to 14 per cent, 2 in 14 on bp35. The normalising baseline moved from the second-best player to the median player per level: from now on a score of 100 per cent for AI means beating every level of every environment at or above median human action efficiency. Human and machine participants received identical information.