o3 on ARC-AGI
The o3 model scored 75.7 per cent on the ARC-AGI contest at an ordinary compute budget and 87.5 at a high one; the previous record was about a third.
Why it matters
A test built deliberately so that scaling could not beat it was beaten by scaling, but of compute at answering time.
ARC-AGI consists of puzzles that require inferring a rule from a few examples; every puzzle is new, and it cannot be memorised from data. Francois Chollet devised it in 2019 precisely as a limit for language models, and it held for five years. o3 broke through it, spending thousands of dollars of compute on a single puzzle in the high mode. Chollet himself called the result a genuine jump but warned it was not general intelligence: problems remained that are easy for a person and out of reach for the model. At that point the model was only announced; public access opened in April of the following year.