OpenAI has initiated a lockdown of its forthcoming Astra model after internal evaluations revealed significant advancements in its agentic coding and cybersecurity capabilities. The company stated it cannot definitively rule out Astra achieving a "critical capability" level for cybersecurity under its Preparedness Framework, which outlines risk assessment and safeguard protocols for advanced AI models.
The Preparedness Framework, established in December 2023, categorizes high-risk areas such as cybersecurity, biological and chemical threats, harmful persuasion, and AI self-improvement. Models undergo safety evaluations prior to deployment, and those deemed to pose unacceptable risks face delayed deployment or the implementation of additional safeguards.
Under this framework, an AI model is classified as possessing critical cybersecurity capabilities if it can autonomously identify previously unknown software vulnerabilities in secure systems or plan and execute sophisticated cyberattacks against well-protected targets with minimal human intervention. OpenAI noted it is applying a similar cautious approach to Astra as it did in 2025 when its AI models began demonstrating advanced biological capabilities.
OpenAI clarified that this is a preliminary assessment and Astra has not yet been formally classified as a critical cybersecurity model. The company is continuing its evaluations to make a final determination. It also explicitly stated that Astra is an upcoming model and was not implicated in any exploitation of Hugging Face.
In response to these findings, OpenAI has bolstered its safeguards and security controls for models exhibiting such advanced capabilities. During Astra's development, stricter security measures were implemented, including isolated testing environments, restricted network and tool access, enhanced protection and encryption of model weights, additional monitoring and detection systems, and sandboxed execution.
Activities involving Astra that do not meet these heightened security requirements have been paused. The company has also introduced comprehensive monitoring to detect risky actions and signs of misalignment within agentic applications.
Prior to Astra's deployment, OpenAI intends to engage government agencies and independent AI safety organizations to evaluate the model's cybersecurity capabilities. External testing partners will also receive guidance and security measures to facilitate the safe evaluation of the model's higher-risk functionalities.






