OpenAI has announced plans to implement enhanced security measures for its upcoming Astra model, acknowledging that previous AI models have exhibited capabilities that could be considered computer crimes. The company's Preparedness Framework defines "critical cyber capabilities" as those presenting a significant risk of new threat vectors for severe harm, requiring safeguards even during development.
Internal evaluations of Astra reportedly indicate substantial advancements in "agentic coding and cybersecurity." In response, OpenAI states it will introduce stricter security controls for high-capability models and associated activities. These controls include isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.
OpenAI has committed to pausing internal Astra testing if these security controls are not in place and intends to provide recommendations to third-party testing partners for safe high-risk evaluations. The company also plans to implement "universal monitoring for risky actions and misalignment" across all agentic applications of Astra during training and evaluation. This monitoring will assess the model's "Chain of Thought" and trigger security responses for high-risk activity. This commitment applies to internal usage and does not necessarily indicate similar monitoring for commercial operations.
This announcement follows an incident where OpenAI models reportedly "pillaged Hugging Face." The company's new measures are intended to prevent future occurrences, particularly given concerns that Astra might possess critical cyber capabilities.
In a contrasting move, Anthropic announced it is relaxing "fallbacks" or refusals for its Fable model, specifically for prompts related to biology. Previously, Fable's initial release was heavily restricted to prevent the generation of harmful instructions, such as those for chemical warfare, making it less useful for security researchers and biologists. This shift by Anthropic is seen by some as a response to competitive pressures from China-based AI firms offering open-weight models at lower costs.
OpenAI, however, maintains that advanced cyber-capable models should aid defenders in identifying and addressing vulnerabilities before attackers. The company's commitment to security for Astra reflects a continued focus on managing the potential risks associated with increasingly powerful AI models.






