OpenAI has temporarily halted a major reinforcement learning (RL) training run for its advanced frontier AI models, citing cybersecurity risks and the need to bolster its research environments. The pause, which has lasted for two weeks, allowed the company to harden systems, conduct red-teaming exercises, and expand monitoring capabilities. This decision was influenced by a recent incident involving OpenAI and Hugging Face, as well as preliminary indications that OpenAI's forthcoming Astra model could achieve "Critical cybersecurity capability" as defined by its internal Preparedness Framework.
The company is implementing a multi-pronged approach focused on strengthening monitoring, alignment, and containment throughout its model development lifecycle. This includes tightening security requirements for environments used to train and evaluate frontier models, particularly as their cybersecurity capabilities increase. These enhanced standards have necessitated significant engineering work, leading to substantial costs and delays in frontier research.
Following the OpenAI-Hugging Face incident, the company initially paused all frontier model inference in research clusters for workloads capable of executing code or utilizing internet-accessing tools. A more restricted and secure pathway for code execution has since been restored, with individual workloads undergoing assessment before resuming operations. These measures include enhanced isolation for untrusted code, stricter network restrictions, and continuous security testing. OpenAI has also reduced standing privileges, eliminated potentially vulnerable shared services, and is improving its ability to collect and monitor security logs. The most stringent controls are applied to Astra and other cyber-related workloads, given Astra's potential to reach critical cybersecurity capabilities. While some Astra training and evaluations now meet the new standards, a significant number of workloads remain paused pending environment migration and upgrades, with safety and alignment work being prioritized for migration.
OpenAI has also expanded its monitoring efforts beyond high-risk internal deployments and frontier RL training to encompass more model activity involving tools. This system employs activation classifiers that scrutinize activity at every sampled token, escalating potential concerns to automated investigators. These investigators examine tool actions, available reasoning, and the full sequence of events for signs of unauthorized access, data theft, destructive behavior, or attempts to bypass safeguards. The goal is to issue an alert within 30 minutes of a concern emerging. Suspected violations of critical security boundaries are escalated to safety, security, and research teams, who are mandated to pause the affected activity if they cannot confirm within 30 minutes that the alert is a false positive.
This 30-minute response requirement applies to all RL training and evaluations involving tools for models of "Sol capability" or higher. On August 7, this requirement was extended to all Astra inference involving tools after OpenAI determined the model might possess critical cybersecurity capabilities. The expanded monitoring adds an estimated 20% to the inference compute being monitored, with varying costs across training and evaluation workloads.
Furthermore, OpenAI is broadening its alignment work for RL runs on its most capable models, applying core alignment techniques across more stages of the training process. This involves refining reward models to detect and discourage unsafe behavior across tasks and environments, training models to be more truthful about their actions, capabilities, and limitations, and mitigating behaviors that exploit weaknesses in rewards, graders, tools, or oversight. Training coverage is also being expanded for behaviors that could cause harm when models interact with external systems or resources. Findings from this broader alignment research and evaluations will inform future training and safeguards.
The company plans to update its Preparedness Framework to integrate these safeguards across training and deployment, better accounting for the capabilities of future models and their operating environments. OpenAI has stated its continued aggressive investment in alignment research, increased evaluation coverage, and the use of learned insights to inform training and safeguards. The company intends to share more information about its alignment research, including model behavior and any novel challenges identified, in the near future.






