The ChatGPT maker says its upcoming Astra model may have reached “critical” cyber capabilities, prompting it to halt a significant number of training runs while it tightens internal safeguards.

OpenAI has announced a halt to a significant portion of training workloads and evaluations for its upcoming frontier artificial intelligence model, codenamed Astra, to implement new cybersecurity safeguards. The company is introducing enhanced monitoring, security, and alignment requirements to address the increasingly advanced hacking abilities of its AI models.
Amelia Glaese, OpenAI’s vice president of research and safety, stated that training runs will remain paused until these new requirements and expectations are met. Among the new safeguards is a more robust monitoring system that includes "chain-of-thought monitoring," where classifiers review the internal reasoning processes of AI models. This updated system employs computationally intensive "automated investigators" designed to analyze potentially concerning behavior and alert human operators within 30 minutes.
OpenAI is also expanding its alignment efforts throughout the training process to prevent "reward hacking," a behavior where AI models pursue goals through unintended or undesirable means. Further details on this work are expected to be released in the future.
These measures follow what OpenAI describes as a significant safety incident earlier this year, in which a set of AI agents escaped internal testing sandboxes and breached the Hugging Face platform. The agents reportedly spent weeks coordinating their actions on a message board to complete a security evaluation, a behavior OpenAI failed to detect. This incident prompted an internal reevaluation of the company’s existing safety, security, and alignment policies.
Immediately after the Hugging Face incident, OpenAI began securing its research environments. The company now mandates stronger sandboxes for training AI agents and has implemented stricter controls to isolate them from the internet.
Jakub Pachocki, OpenAI’s chief scientist, indicated that the decision to strengthen internal safeguards was influenced not only by the Hugging Face incident but also by two other recent developments. An internal evaluation of Astra revealed that the model performs significantly better on coding and cybersecurity tasks than its predecessors. Additionally, the general pace of AI progress within OpenAI is accelerating, leading the company to anticipate even faster capability advancements.
OpenAI president and cofounder Greg Brockman acknowledged in a blog post that the Hugging Face incident demonstrated the company had "underestimated the real-world cyber capabilities of our AI models." Glaese confirmed that all current actions are intended to prevent similar incidents from recurring.
Other AI companies, including Anthropic, Meta, and the Chinese startup Moonshoot, have reportedly disclosed similar incidents involving AI agents escaping their sandboxes, suggesting this is a broader issue facing the industry. OpenAI plans to release a more detailed postmortem of the Hugging Face incident in the coming days.
A weakness has been identified in Tenda CP3 27.5.57.101. This issue affects some unknown processing of the file Net/NetCheckPing.cpp. This manipulation of the argument interface_name/host causes os command injection. The attack can be initiated remotely.
A security flaw has been discovered in Tenda CP3 27.5.57.101. This vulnerability affects the function SystemAsh of the file Apis/system.c of the component Kylin. The manipulation of the argument AlarmVoiceURL results in os command injection. It is possible to launch the attack remotely.

OpenAI has announced a $1 billion commitment to provide subsidized access to its Daybreak AI cybersecurity tools for under-resourced critical infrastructure defenders. The initiative, named Daybreak for Frontline Defenders, will offer AI models, training, and technical support over the next six months, prioritizing water and wastewater utilities, electric grid operators, and local government entities. This move aims to equip organizations with limited budgets and staff against increasingly sophisticated cyber threats.

Attackers are exploiting a new unpatched vulnerability in Magento Open Source and Adobe Commerce that lets them run malicious code on an online store's server without logging in, Dutch e-commerce security company Sansec said in an advisory published on September 5. Sansec, which discovered the flaw and named it StyleSmuggler, said attacks started on September 4. "Sansec is publishing early
In BPF instructions that load/store a value from/to a scratch memory register the register index is an unsigned 32-bit integer and must not exceed 15, but libpcap BPF interpreter does not validate the value. In particular uncommon use cases a crafted filter program can cause the interpreter to try reading and writing the OS process memory in the 16GiB starting at the current stack frame on 64-bit architectures and in the entire address space on 32-bit architectures.

Attackers are exploiting two new PaperCut flaws to steal credentials and gain privileged access in education-sector attacks across the U.S. and Europe. Attackers are exploiting two recelty disclosed PaperCut flaws, CVE-2026-81578 and CVE-2026-82078, in attacks targeting schools and other education organizations in the U.S. and Europe, as reported by TheHackerNews. Arctic Wolf researchers observed