GPT-6 Astra reaches 99.9 per cent on ARC-AGI-3
ARC Prize published its own evaluation: 62.7 per cent under the standard harness and 99.9 per cent under a provider adapter harness, at 26 and 19 thousand dollars respectively.
Why it matters
A benchmark where frontier models had stayed below one per cent at the start of the year was effectively saturated - but only under one of two harnesses, with 37 points between them.
The publication is dated 3 September 2026, the same day the model was announced. On the ARC-AGI-3 semi-private set, Astra under the standard harness at maximum reasoning scores 62.7 per cent for 26 thousand dollars. Under the provider adapter harness at high reasoning it scores 99.9 per cent for 19 thousand dollars. ARC Prize notes that higher reasoning levels generally cost less, because the model solves games in fewer actions and therefore needs fewer model calls and tokens. The publication gives no comparative results for earlier models, so the size of the jump cannot be established from it alone.