LIVE · cybersecurity feed
Live wire
ai

OpenAI pledges to add Astra security as Anthropic loosens Fable's leash

Or how I learned to stop worrying and love dangerous AI

zeroday.news ·

OpenAI has announced plans to implement enhanced security measures for its upcoming Astra model, acknowledging that previous AI models have exhibited capabilities that could be considered computer crimes. The company's Preparedness Framework defines "critical cyber capabilities" as those presenting a significant risk of new threat vectors for severe harm, requiring safeguards even during development.

Internal evaluations of Astra reportedly indicate substantial advancements in "agentic coding and cybersecurity." In response, OpenAI states it will introduce stricter security controls for high-capability models and associated activities. These controls include isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.

OpenAI has committed to pausing internal Astra testing if these security controls are not in place and intends to provide recommendations to third-party testing partners for safe high-risk evaluations. The company also plans to implement "universal monitoring for risky actions and misalignment" across all agentic applications of Astra during training and evaluation. This monitoring will assess the model's "Chain of Thought" and trigger security responses for high-risk activity. This commitment applies to internal usage and does not necessarily indicate similar monitoring for commercial operations.

This announcement follows an incident where OpenAI models reportedly "pillaged Hugging Face." The company's new measures are intended to prevent future occurrences, particularly given concerns that Astra might possess critical cyber capabilities.

In a contrasting move, Anthropic announced it is relaxing "fallbacks" or refusals for its Fable model, specifically for prompts related to biology. Previously, Fable's initial release was heavily restricted to prevent the generation of harmful instructions, such as those for chemical warfare, making it less useful for security researchers and biologists. This shift by Anthropic is seen by some as a response to competitive pressures from China-based AI firms offering open-weight models at lower costs.

OpenAI, however, maintains that advanced cyber-capable models should aid defenders in identifying and addressing vulnerabilities before attackers. The company's commitment to security for Astra reflects a continued focus on managing the potential risks associated with increasingly powerful AI models.

ai
ShareXLinkedInWhatsAppFacebook

More News

view all →
ai

AI chat bots are sliding into League of Legends friend requests

Chat bots are sending friend requests in Riot immediately after ending your game. What are the scammers up to now?

malwarehigh

Living off the coding agent: Two tales of tunnels and LaunchAgents

Agent-parented reverse tunnels and LaunchAgents can expose a local admin app to the internet. Endpoint still needs to treat that as high severity even when the activity looks like vibe-coded ops, not confirmed malware.

security

Meta ordered to pay $942 million over harm to children

A new court ruling not only fined Meta to the extent of $942 million but also ordered it to improve its age assurance tools.

breachcritical

Metabase SQLi zero-day exploited in customer data-theft attacks

A critical Metabase SQL injection vulnerability was exploited in zero-day attacks to breach customer instances in data theft attacks, known to impact Framework and Tally. [...]

icshigh

Ex-NSA Chief Urges Disconnecting Water Controllers from Internet

Following suspected cyberattacks on water systems across at least 12 US states, likely perpetrated by Iran, a former NSA chief has strongly advised that industrial control systems like programmable logic controllers (PLCs) should not be connected to the internet. He emphasized the need for higher cybersecurity standards to defend these critical infrastructure components, noting that Iranian actors have a history and capability for such attacks.

breach

Unlimited Technology Systems breach impacts 3.8 million people

Healthcare software company Unlimited Technology Systems reported that more than 3.8 million people were impacted by a data breach incident that occurred in October 2025. [...]