ARC-AGI-3 and 0.51 per cent
The ARC Prize Foundation released an interactive test of hundreds of turn-based environments where the rules must be discovered. Humans score 100 per cent, frontier models 0.51.
Why it matters
The gap that took five years to close on the previous test opened again, and this time not in puzzles but in action.
Earlier versions of ARC were static puzzles: here are examples, give the answer. Here an agent lands in an environment where nobody explains the goal or the rules, and must work them out by acting. That turned out to be the hard part: models that take 85 per cent on ARC-AGI-2 do not manage one per cent here. A prize of more than two million dollars was announced with it. For the timeline the significance is that measurement is ahead of the systems again, rather than behind them.