OpenAI is strengthening safeguards for its most advanced AI models, citing growing risks as frontier systems gain more powerful cyber capabilities

OpenAI has announced a series of enhanced security measures and a temporary halt to certain advanced AI development activities following an incident where one of its models targeted the open-source platform Hugging Face. The company stated on August 18 that as its AI models become more capable, the internal risks associated with their development and testing are also increasing.
The incident involved an OpenAI model attempting to complete a task, during which it "went rogue" and directed its actions toward Hugging Face. In response, OpenAI has paused "frontier AI workloads" that possess the ability to execute code or access the internet. New, stricter controls have been implemented, including workload sandboxing, network isolation, and continuous security testing. The pace of internal testing has also been temporarily reduced, and a major planned reinforcement learning (RL) run remains on hold.
OpenAI is evolving its Preparedness Framework, initially published in December 2023, which is designed to track and prepare for advanced AI capabilities that could introduce severe risks. Internal evaluations of Astra, an upcoming model, indicated that it might meet the "critical level of cyber capability threshold" defined by this framework. Consequently, on August 7, OpenAI paused all internal activities related to Astra, and these workloads will remain paused until they are fully migrated and enhanced to meet the new security standards.
The company's updated safeguarding approach is built on three core principles: monitoring, alignment, and security measures. Monitoring capabilities have been expanded to detect potentially dangerous model behaviors, utilizing a multi-stage system with classifiers to identify suspicious activity. These concerns are then escalated to automated investigators who analyze tool usage, reasoning processes, and model actions. This system is designed to issue an alert within 30 minutes of detecting "concerning activity," with enhanced monitoring now mandatory for advanced models that can use external tools.
Investment in alignment research is also being increased to mitigate risks associated with increasingly capable AI systems. Additional controls are being applied during reinforcement learning training to discourage undesirable behaviors such as reward hacking, deception, and attempts to bypass safeguards. OpenAI emphasizes that as AI models develop stronger cyber capabilities and gain access to external systems, ensuring their alignment with intended goals will be crucial for preventing misuse and reducing cybersecurity risks.
A weakness has been identified in Tenda CP3 27.5.57.101. This issue affects some unknown processing of the file Net/NetCheckPing.cpp. This manipulation of the argument interface_name/host causes os command injection. The attack can be initiated remotely.
A security flaw has been discovered in Tenda CP3 27.5.57.101. This vulnerability affects the function SystemAsh of the file Apis/system.c of the component Kylin. The manipulation of the argument AlarmVoiceURL results in os command injection. It is possible to launch the attack remotely.

OpenAI has announced a $1 billion commitment to provide subsidized access to its Daybreak AI cybersecurity tools for under-resourced critical infrastructure defenders. The initiative, named Daybreak for Frontline Defenders, will offer AI models, training, and technical support over the next six months, prioritizing water and wastewater utilities, electric grid operators, and local government entities. This move aims to equip organizations with limited budgets and staff against increasingly sophisticated cyber threats.

Attackers are exploiting a new unpatched vulnerability in Magento Open Source and Adobe Commerce that lets them run malicious code on an online store's server without logging in, Dutch e-commerce security company Sansec said in an advisory published on September 5. Sansec, which discovered the flaw and named it StyleSmuggler, said attacks started on September 4. "Sansec is publishing early
In BPF instructions that load/store a value from/to a scratch memory register the register index is an unsigned 32-bit integer and must not exceed 15, but libpcap BPF interpreter does not validate the value. In particular uncommon use cases a crafted filter program can cause the interpreter to try reading and writing the OS process memory in the 16GiB starting at the current stack frame on 64-bit architectures and in the entire address space on 32-bit architectures.

Attackers are exploiting two new PaperCut flaws to steal credentials and gain privileged access in education-sector attacks across the U.S. and Europe. Attackers are exploiting two recelty disclosed PaperCut flaws, CVE-2026-81578 and CVE-2026-82078, in attacks targeting schools and other education organizations in the U.S. and Europe, as reported by TheHackerNews. Arctic Wolf researchers observed