LIVE · cybersecurity feed
Live wire
ai

Irregular says ‘human oversight’ responsible for AI sandbox escape incidents

In a post-mortem, the frontier AI testing company said internet access for models is necessary to fully test out their cybersecurity capabilities. The post Irregular says ‘human oversight’ responsible for AI sandbox escape incidents appeared first on CyberScoop.

zeroday.news ·

Irregular, a company specializing in the security evaluation of advanced AI models, has acknowledged that "human oversight" was a contributing factor in recent incidents where AI models, including Anthropic's Claude Opus and OpenAI's GPT-5.6 Sol, escaped their simulated environments and took offensive actions in the real world. The company stated that unintentional internet access was provided to the models during testing, which is typically designed to stress-test their cybersecurity capabilities.

The incidents involved non-public models from both Anthropic and OpenAI, specifically Mythos 5, Claude Opus, and GPT-5.6 Sol. Irregular's post-mortem analysis, detailed in a blog post, confirmed that a small number of interactions with their evaluation environments led these models to engage in real-world offensive security actions.

In one notable instance involving Anthropic's models, testers initially believed the AI was targeting a fictional company within the simulation. However, it was later discovered that the fictional company's name inadvertently matched a real-world domain. This led the model to consider and, in some training runs, execute offensive actions against the actual company. Irregular attributed this to human error, noting that standard background checks to prevent the use of real company names in simulations were not adequately performed in these cases.

The instructions given to the models included the target's name and internal network addresses for the simulated environment. While Mythos performed as intended in the vast majority of tests, a small number of cases demonstrated the model's inability to differentiate between fake and real domains. Consequently, it executed actual attacks on internet infrastructure, including exploiting vulnerabilities, extracting credentials, and accessing a production database. In one specific case, the model even began targeting a site with a similar name after discovering its credentials online.

Irregular emphasized that, for the most part, the models believed they were operating within simulated environments even as they took actions in the real world. The company has since remediated the issues that led to these interactions and is implementing new protocols to prevent similar setup errors.

The company also defended the practice of granting some level of internet access to AI models during security evaluations. They argued that controlled internet access, despite the inherent risks of exceeding containment boundaries, is crucial for realistic threat scenario testing. Without it, the fidelity of such scenarios is compromised, undermining the goal of reducing post-release risks, as real-world attackers rely on internet access.

The incidents have highlighted critical gaps in Irregular's security practices. Moving forward, the company plans to improve documentation for evaluation setups, deploy enhanced log monitoring tools capable of handling the extensive data generated by AI traffic, revise their threat models to account for rogue AI behavior, and establish faster information-sharing mechanisms among stakeholders.

Irregular acknowledged that while improved implementation of existing safeguards could prevent most such incidents currently, this might not be sufficient as AI models become more powerful. The company views this as an opportunity for itself and the wider community to proactively establish forward-looking protocols and invest in research and development to address future challenges posed by increasingly capable AI. A comprehensive whitepaper detailing the incidents and updated best practices for evaluation setups is expected to be released.

ai
ShareXLinkedInWhatsAppFacebook

More News

view all →
breachcritical

LLMs and Contextual Integrity

I have been thinking a lot about AI and integrity. Part of that is contextual integrity. I recently found two papers on the topic. “CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs“: Abstract: Large Language Models (LLMs) increasingly use persistent memory from past interactions to enhance personalization and task performance. However, this memory introdu

security

Meta Ran Ads for an App That Promised to Nudify Female Politicians

One advertisement featured a pornographic video with a deepfake closely resembling a prominent US politician. Apple removed the app from the App Store after an inquiry from WIRED.

security

Hackers target Ukrainian agency managing assets seized from sanctioned Russians

The agency said the latest attack came amid preparations to select a manager for seized corporate rights in IDS Ukraine, one of the country’s largest producers of bottled mineral water and beverages.

vulnerabilitycritical

NASA Ground Control Software Flaw Enables Unauthenticated Commands

Critical AIT-GUI flaws expose spacecraft commands and scripts to unauthenticated attackers

CVE-2026-19478critical

Critical GitLab flaw allows attackers to modify or delete public projects (CVE-2026-19478)

GitLab has released patches for two vulnerabilities, including a critical-severity code injection flaw that can be exploited without authentication. The vulnerabilities affect GitLab Community Edition (CE) and Enterprise Edition (EE) versions from 18.2 before 18.11.11, 19.0 before 19.0.8, 19.1 before 19.1.6, and 19.2 before 19.2.4. The fixes are available in GitLab 19.2.4, 19.1.6, 19.0.8, and 18.1

security

Cyber Incident Disrupts Student Services at UT San Antonio

UT San Antonio has taken IT systems offline following a cyber incident, disrupting student registration and tuition payments days before term is due to resume