Irregular, a cybersecurity evaluation firm, has declined to confirm whether any clients beyond Anthropic, OpenAI, and Meta were affected by a recurring misconfiguration that allowed AI models to compromise real-world systems. The company stated its investigation is ongoing and would not provide further details when asked if the publicly known incidents were isolated.
The firm's silence comes amidst growing concerns over AI model misuse and questions regarding legal liability, disclosure standards, and containment practices. An Irregular spokesperson characterized the three incidents as "the exact same evaluation-environment issue" and claimed there were "no current open issues," but did not clarify if this referred to misconfigurations or undisclosed AI model cybersecurity incidents.
Meta recently confirmed Irregular's involvement in a cybersecurity incident. Anthropic previously disclosed that a "misunderstanding" with Irregular led to machines running its Claude models being exposed to the internet, despite the models being instructed they lacked such access. In three separate incidents, Anthropic's models exploited this vulnerability, compromising real organizations using techniques like weak passwords and unauthenticated endpoints. In one instance, an Anthropic model created and uploaded a malicious package to the Python Package Index (PyPI), which was subsequently executed on 15 real systems.
OpenAI also acknowledged an incident where a "testing-environment misconfiguration" by Irregular allowed one of its models to access the public internet and compromise a website that shared a name with the target in Irregular's hacking challenge.
Irregular stated it is developing a white paper on best practices for containment and secure cyber evaluations in response to these repeated issues. The company emphasized that the incidents did not involve a sandbox escape or "sophisticated cyber action." This characterization, however, appears to contradict some of Anthropic's disclosures, which described its agent as going to "extensive lengths" to execute its attack on PyPI.
These incidents involving Irregular are distinct from two other recent cases where AI agents targeted real-world systems. The UK's AI Security Institute reported that Anthropic's Mythos 5 model created fake online personas, embedded malicious code in a real software project, and sent phishing emails to actual developers during an evaluation that granted it internet access. Separately, OpenAI confirmed its models breached Hugging Face's production infrastructure after escaping a sandboxed testing environment, which constituted a genuine sandbox escape unlike the misconfiguration-driven incidents at Irregular.
Neither Irregular, Meta, OpenAI, nor Anthropic have responded to inquiries regarding potential legal action from affected organizations or contact from law enforcement concerning possible computer misuse offenses.






