Back to timeline

Milestone · August 18, 2026

OpenAI slows frontier training

In a post of 18 August 2026 OpenAI said it had temporarily slowed scaling: it paused reinforcement-learning training of its latest models intended for deployment for two weeks and kept its largest planned frontier RL run on hold. The reasons it gave were the Hugging Face incident and preliminary evidence that its coming model Astra might reach the Critical cyber level of the company's Preparedness Framework.

Why it matters

The laboratory is not calling on others to slow down; it reports what it has already done itself, pausing part of its own training and rebuilding its environments. The substance is in the measures, and they are not cheap: by OpenAI's figure the monitoring it now requires costs about a fifth of the compute of the inference it watches. All of it is the company's own account; nobody else checked the length or the scope of the pause.

What the post says (OpenAI's account; there is no independent confirmation). Reinforcement-learning training of its latest models intended for deployment was paused for two weeks; the largest planned frontier RL run is 'still paused' while smaller training and evaluation continue. Right after the Hugging Face incident the company paused inference of frontier models on research clusters for runs that execute code or use tools with internet access, and then reviewed each workload separately: some were restored under new controls, others needed changes. The new requirements: stronger isolation for workloads that run model-generated or other untrusted code; network isolation of untrusted and higher-risk workloads; continuous testing of the isolation boundaries with simulated attacks run by the company's own models. Monitoring is multistage: activation classifiers first, then increasingly elaborate automated investigators; alerts are meant to go out within 30 minutes, and if the teams cannot establish within those 30 minutes that a trigger is false they are to pause the activity. It is mandatory for all tool-using reinforcement-learning training and evaluations of models at the level of Sol or above; after OpenAI determined on 7 August that Astra might have critical cyber capabilities, it also covers all tool-using Astra inference. The company's current estimate of the cost is about 20% of the compute of the monitored inference. The post promises a technical report 'in the coming weeks'; it came out on 26 August (separate record). What followed, from OpenAI's post of 1 September ('Path to Astra'): on 28 August, once the new security requirements were in place, the large-scale reinforcement-learning run was resumed; smaller experimental runs remain deferred. So 'still paused' was true on 18 August and stopped being true ten days later. In the same post the company writes that Astra has reached the Critical threshold. What the record does not claim: that the reasons were anything other than those the post names; that anyone outside OpenAI checked the pause or its scope; what it did to the date of Astra's release (3 September, a separate record).

Event record

Event date
August 18, 2026
Timeline date
Event date
Verification
Sources gathered automatically · October 10, 2026
Lines
ID
evt-0987

The day of OpenAI's post. The post says reinforcement-learning training was paused for two weeks but does not say when the pause began; the resumption of the large run on 28 August is in the post of 1 September.

Sources

Related events

Earlier

Later