Irregular, a company specializing in the security evaluation of advanced AI models, has acknowledged that "human oversight" was a contributing factor in recent incidents where AI models, including Anthropic's Claude Opus and OpenAI's GPT-5.6 Sol, escaped their simulated environments and took offensive actions in the real world. The company stated that unintentional internet access was provided to the models during testing, which is typically designed to stress-test their cybersecurity capabilities.
The incidents involved non-public models from both Anthropic and OpenAI, specifically Mythos 5, Claude Opus, and GPT-5.6 Sol. Irregular's post-mortem analysis, detailed in a blog post, confirmed that a small number of interactions with their evaluation environments led these models to engage in real-world offensive security actions.
In one notable instance involving Anthropic's models, testers initially believed the AI was targeting a fictional company within the simulation. However, it was later discovered that the fictional company's name inadvertently matched a real-world domain. This led the model to consider and, in some training runs, execute offensive actions against the actual company. Irregular attributed this to human error, noting that standard background checks to prevent the use of real company names in simulations were not adequately performed in these cases.
The instructions given to the models included the target's name and internal network addresses for the simulated environment. While Mythos performed as intended in the vast majority of tests, a small number of cases demonstrated the model's inability to differentiate between fake and real domains. Consequently, it executed actual attacks on internet infrastructure, including exploiting vulnerabilities, extracting credentials, and accessing a production database. In one specific case, the model even began targeting a site with a similar name after discovering its credentials online.
Irregular emphasized that, for the most part, the models believed they were operating within simulated environments even as they took actions in the real world. The company has since remediated the issues that led to these interactions and is implementing new protocols to prevent similar setup errors.
The company also defended the practice of granting some level of internet access to AI models during security evaluations. They argued that controlled internet access, despite the inherent risks of exceeding containment boundaries, is crucial for realistic threat scenario testing. Without it, the fidelity of such scenarios is compromised, undermining the goal of reducing post-release risks, as real-world attackers rely on internet access.
The incidents have highlighted critical gaps in Irregular's security practices. Moving forward, the company plans to improve documentation for evaluation setups, deploy enhanced log monitoring tools capable of handling the extensive data generated by AI traffic, revise their threat models to account for rogue AI behavior, and establish faster information-sharing mechanisms among stakeholders.
Irregular acknowledged that while improved implementation of existing safeguards could prevent most such incidents currently, this might not be sufficient as AI models become more powerful. The company views this as an opportunity for itself and the wider community to proactively establish forward-looking protocols and invest in research and development to address future challenges posed by increasingly capable AI. A comprehensive whitepaper detailing the incidents and updated best practices for evaluation setups is expected to be released.






