LIVE · cybersecurity feed
Live wire
Critical Zimbra RCE flaw now actively exploited in attacksExploitation Expected for Critical Authentication Bypass Patched in Citrix NetScalerCVE-2026-19478 · Critical GitLab Flaw Exploited Shortly After DisclosureCVE-2026-32475 · Elementor Pro Flaw Could Let Unauthenticated Attackers Upload PHP and Execute Code8,539 reasons to rethink how vulnerabilities get patched'Not a theoretical risk,' feds warn as attackers use AI-made code to hack critical infrastructure controllersNSA, FBI warns of hackers using AI-generated tools in attacks on critical infrastructure technologyUS warns of AI-powered attacks on Siemens PLCs in critical infrastructureCVE-2024-39943 · Operation CameraSwarm Compromised 14,000+ Dahua CamerasCVE-2026-19490 · CVE-2026-19490: Critical Vulnerability Affecting Citrix NetScaler ADC and NetScaler Gateway
ai

OpenAI Overhauls Model Security With Sandboxing, 30-Minute Alerts, and Training Pauses

The action taken by OpenAI comes in light of the Hugging Face incident and the discovery of the Astra model’s advanced capabilities. The post OpenAI Overhauls Model Security With Sandboxing, 30-Minute Alerts, and Training Pauses appeared first on SecurityWeek.

zeroday.news ·

OpenAI has reportedly implemented significant overhauls to its model security protocols, introducing sandboxing for models, a 30-minute alert system for potential security incidents, and the capability to pause model training. These changes are understood to be a direct response to a recent incident involving Hugging Face and the observed advanced capabilities of a model identified as "Astra."

The introduction of sandboxing aims to isolate models during development and deployment, thereby containing potential vulnerabilities or malicious behaviors. In a sandboxed environment, a model's access to system resources and external networks is restricted, limiting the damage it could inflict if compromised or if it exhibits unintended emergent properties. This is a common security practice in software development, now being applied more rigorously to AI model lifecycles.

The new 30-minute alert system is designed to provide rapid notification of suspicious activities or security anomalies. Such a short response window suggests a focus on minimizing the dwell time of threats and enabling quick intervention. This capability likely leverages real-time monitoring and anomaly detection systems that flag deviations from normal model behavior or operational parameters.

Furthermore, the ability to pause model training represents a critical control mechanism. In scenarios where a model might be exhibiting undesirable or potentially dangerous behaviors during its training phase, or if a security vulnerability is discovered in the training pipeline or data, halting the process can prevent further propagation of issues. This allows for investigation and remediation before the model is deployed or further developed.

The reported catalyst for these changes includes an unspecified incident involving Hugging Face. While details of this incident are not provided, it likely highlighted specific vulnerabilities or attack vectors relevant to large language models or their development environments. Similarly, the "discovery of the Astra model’s advanced capabilities" suggests that internal or external assessments revealed emergent properties or potential risks that necessitated a re-evaluation of existing security measures.

These security enhancements are indicative of a broader industry trend towards more robust and proactive security postures for advanced AI systems. As AI models become more complex and integrated into critical applications, the potential impact of security breaches or unintended model behaviors increases. Implementing controls like sandboxing, rapid alerting, and training pauses are essential steps in managing the unique risks associated with sophisticated AI development and deployment.

Such measures reflect a growing recognition within the AI community that model security extends beyond traditional software security, encompassing the integrity of training data, the safety of model outputs, and the control over emergent model behaviors. This comprehensive approach is becoming standard practice as AI systems move from research environments to widespread application.

ai
ShareXLinkedInWhatsAppFacebook

More News

view all →
ai

Managing the cyber risk of agentic AI

Use safeguards, sandboxing and active oversight to realise the benefits of autonomous systems while limiting the unintended activity.

malware

New Manic Android malware can exfiltrate data through nearby devices

A new Android malware named Manic targeting users in multiple European countries has a fallback data exfiltration mechanism that uses nearby infected devices. [...]

security

Police Are Hiding Their Use of Flock Surveillance Cameras

A usage policy for Flock license plate reader cameras tells police not to talk about the cameras: When cops use Flock to arrest someone in Wapello County, Iowa, they don’t want them to know. A usage policy for the automated license plate reader cameras in the county tells police, in no uncertain terms, to keep them a secret: “DO NOT MENTION ALPR USAGE TO THE OCCUPANTS OF THE VEHICLE,” the policy d

vulnerabilitycritical

Critical Zimbra RCE flaw now actively exploited in attacks

CERT Polska, the Polish Computer Emergency Response Team (CERT), warned that attackers have begun exploiting a critical vulnerability in Zimbra Collaboration Suite (ZCS). [...]

phishing

Def Con Attendees Targeted by Persistent Phishing Campaign

Huntress researcher explains how they were targeted by an elaborate and persistent phishing scam following Def Con

security

US Indicts 17 Iranians Over Years-Long Cyber Espionage Campaign

The US charged 17 Iranians over a years-long hacking campaign that stole 31TB from universities, companies and government agencies worldwide. Eight years after the original indictment first went public, US prosecutors just added eight more names to the list. The Justice Department unsealed a superseding indictment this week charging 17 members of the Mabna Institute, […]