OpenAI says it has temporarily slowed parts of its frontier-model development work while it raises the security, monitoring and alignment standards around highly capable systems. In a new update, the company says preliminary evidence suggests that an upcoming model, Astra, may meet the “Critical” cybersecurity capability threshold in its Preparedness Framework.
The announcement is significant because it describes a concrete change to research operations, rather than only a future policy commitment. OpenAI says it paused reinforcement-learning training on its latest models intended for deployment for two weeks, and that its largest planned frontier reinforcement-learning run remains on hold while it conducts smaller training runs and evaluations.
Why OpenAI changed its pace
OpenAI links the decision to two recent developments: an OpenAI–Hugging Face incident and signals from its internal research on advanced cyber capabilities. Its conclusion is that the standards used to monitor, align and secure models need to stay ahead of the risks created as those models become more capable and gain access to tools or complex environments.
The company is not presenting the pause as a general halt to all work. It says some research workloads have already resumed under more limited controls, while other workloads need further changes before they can restart. The stated goal is to obtain more evidence about model behaviour, validate safeguards and improve the alignment evidence before proceeding with the largest run.
For people watching the practical governance of frontier AI, this is an unusually operational disclosure. The company is acknowledging that safeguards can impose cost and delay, and that they can affect the sequencing of model development. It also puts more detail behind the often broad terms “security” and “alignment”.
Three layers of safeguards
OpenAI describes its approach as three reinforcing layers: monitoring, alignment and security measures. Monitoring is intended to detect concerning behaviour and support a response. Alignment is intended to reduce the likelihood of harmful or unauthorised actions. Security measures limit what systems can access or affect. The company says all three need to scale alongside model capability.
In research environments, the changes include stronger workload isolation for systems that execute model-generated or otherwise untrusted code. OpenAI also says it has added controls to separate higher-risk workloads from the internet and other internal networks, with the aim that a compromise in one workload or service cannot by itself create broad unauthorised access.
It is also increasing continuous security testing. The update describes removing potentially vulnerable shared services, reducing standing privileges, improving logging and using models to test security boundaries against simulated attacks. OpenAI says the strictest safeguards now apply to workloads involving Astra or cyber models, although some of those workloads remain paused until they meet the new standard.
Expanded monitoring for tool use
Another major element is a broader approach to chain-of-thought monitoring. OpenAI says its new setup begins with classifiers that inspect a model’s internal activity at every sampled token. Potential concerns are escalated to more capable automated investigators, which can examine tool actions, available reasoning and the sequence of activity for signs of unauthorised access, data theft, destructive behaviour or attempts to defeat safeguards.
The company says it aims to alert within 30 minutes after concerning activity is surfaced. Where a likely critical security-boundary violation is identified, safety, security and research teams are paged. If they cannot determine within that period that the signal is a false positive, the activity is expected to be paused.
OpenAI says the monitoring requirement now applies to reinforcement-learning training and evaluations involving tools for models of Sol capability or higher. After determining that Astra may have critical cyber capabilities on 7 August, it added the requirement for Astra inference with tools as well. The company estimates that monitoring currently adds roughly 20% to the inference compute being monitored, while noting that the cost varies by workload.
Alignment remains part of the response
The update also connects security controls to alignment research. OpenAI says it is applying core alignment techniques across more stages of training for its most capable reinforcement-learning runs. This includes improving reward models, encouraging models to be candid about actions and limits, and reducing behaviours that exploit weaknesses in rewards, graders, tools or oversight.
That connection is important: the company is treating cyber risk as more than an access-control problem. Better isolation can reduce the impact of a failure, while monitoring can help detect it, but both are strongest when paired with training intended to reduce unsafe behaviour in the first place.
What happens next
OpenAI says it will evolve its Preparedness Framework to bring these safeguards together across training and deployment. It also plans further work on model-assisted security, monitoring and alignment research, and says it intends to involve external organisations and share more as the approach develops.
The announcement does not promise a timetable for the held frontier run or disclose full technical details of Astra. Instead, it draws a clear line: when evidence suggests a model has reached a higher level of cyber capability, the research environment and controls must reach a higher standard too. That is a consequential precedent for how OpenAI describes balancing speed, capability and safety in the next phase of its model development.