An independent measure ranks GPT-5.5 first
On 23 April 2026, the same day OpenAI released GPT-5.5, Artificial Analysis - which had pre-release access to the model at all five reasoning-effort levels - published its own measurement: the model topped the Artificial Analysis Intelligence Index by 3 points, breaking a three-way tie with Anthropic and Google, and recorded the highest accuracy on the AA-Omniscience knowledge benchmark (57%), alongside the highest hallucination rate among the compared models (86%).
Why it matters
An independent measurer confirmed the manufacturer's model led on several tests on the very day it launched, and at the same time named a trade-off the company's own announcement does not: the highest knowledge accuracy comes with the highest rate of invented answers among the compared models. An editorial assessment.
GPT-5.5 (xhigh effort) led five headline Artificial Analysis evaluations (including Terminal-Bench Hard, GDPval-AA and the newly hosted APEX-Agents-AA), trailing only other OpenAI models on CritPt and AA-LCR and coming second to Gemini 3.1 Pro Preview on three further evaluations. On GDPval-AA (real-world economically valuable tasks) the model scored an Elo of 1785, ahead of Claude Opus 4.7 (max) by about 30 points and Gemini 3.1 Pro Preview by about 470. On AA-Omniscience (knowledge and hallucination) its accuracy of 57% is the highest among the compared models and 14 points above GPT-5.4, but its hallucination rate is 86%, versus 36% for Opus 4.7 (max) and 50% for Gemini 3.1 Pro Preview; the accuracy gain came mostly from added knowledge, not from a comparable drop in the tendency to answer when unsure. At the medium effort level the model matches Opus 4.7 (max) on the Intelligence Index at a quarter of the cost (about $1,200 versus $4,800), though Gemini 3.1 Pro Preview reaches the same score for $900; per-token pricing has doubled since GPT-5.4, but roughly a 40% cut in token use yields a net ~20% increase in the cost of running the Index. What the record does not claim. The source is a measurer, not a peer-reviewed publication; the comparisons rest on Artificial Analysis's own evaluations, which the company does not fully disclose. The record does not claim that a higher hallucination rate makes the model worse overall - only that, on this test, the gain in accuracy was not matched by a comparable drop in the hallucination rate.