Back to timeline

Availability · February 12, 2026

Codex-Spark on Cerebras chips

On 12 February 2026 OpenAI released GPT-5.3-Codex-Spark as a research preview: a smaller model for real-time coding that on Cerebras hardware delivers more than 1000 tokens per second.

Why it matters

An OpenAI coding model was served on specialised non-Nvidia hardware at a speed that lets code be edited with almost no delay. An editorial assessment: access is narrow, ChatGPT Pro subscribers and a few API partners only.

What the page says. A research preview for ChatGPT Pro users in the Codex app, CLI and VS Code extension; in the API for a small set of partners. 128k-token context, text only. The model runs on Cerebras' Wafer Scale Engine 3; per the page it is the first milestone of the partnership announced in January. Separately OpenAI reports that a persistent WebSocket connection and a reworked stack cut per-roundtrip overhead by 80%, per-token overhead by 30% and time to first token by 50%; that path is to become the default for all models. According to OpenAI the model has no plausible chance of reaching the "high" capability threshold for cybersecurity or biology under its preparedness framework. What the record does not claim. Results on SWE-Bench Pro and Terminal-Bench 2.0 are a chart with no numbers in the text; they are the developer's and are not entered. No independent speed measurement was found.

Event record

Event date
February 12, 2026
Timeline date
Event date
Verification
Sources gathered automatically · September 29, 2026
Lines
ID
evt-0952

Sources

Related events

Antecedents for this event are still being researched.