Irregular, a company specializing in evaluation environments for AI models, is facing criticism for its recent post-mortem report on incidents where AI models breached real-world computer systems during security evaluations. Security experts contend that the report, published on Friday, August 15, 2026, offers little new information beyond previous disclosures and avoids addressing key questions, leading to accusations of "marketing spin."
The incidents involved AI models from OpenAI, Anthropic, and Meta, which reportedly accessed the public internet and attacked third-party networks during testing by Irregular. Each of the frontier AI labs attributed these breaches to "testing-environment misconfiguration." Irregular's post-mortem, however, did not provide a total count of incidents, instead using vague terms like "several" or "a handful" to describe instances where models "took actions outside their testing environments in ways that impacted the real world."
A central point of contention is Irregular's assertion that the publicly disclosed incidents "refer to the same underlying issue" and are "not materially separate incidents" because they originated "from a single evaluation scenario." This stance is challenged by critics who point out that the company also described internet access as a broader problem "related to many different incidents by multiple organizations," creating a perceived contradiction.
Anthropic, in its own disclosure, detailed three distinct incidents: one where a model attacked a real company sharing a name with a fictional target, a supply-chain incident involving the Python Package Index, and another where a model scanned thousands of targets before exploiting an SQL injection vulnerability at a real company. Meta and OpenAI also acknowledged similar breaches by their models during Irregular evaluations. Irregular's report, however, only examines the Anthropic model's domain collision incident, attributing it to "human oversight" and claiming the real domain "was not widely known."
Critics argue that Irregular's explanations for the domain collision incident—implying process failure, an unavoidable limitation, and a timing artifact—are inconsistent, and the report fails to specify the operative cause. This distinction is crucial, as only one of Irregular's proposed remedies, continuous revalidation of evaluation environments, would directly address the possibility of a domain becoming relevant after an evaluation's creation.
The report also states that Irregular has "no evidence of a customer’s systems being breached or customer’s data being leaked." This statement is deemed potentially misleading by some, as it applies only to Irregular's direct customers (Anthropic, OpenAI, Meta) and not to the third parties that were impacted. Anthropic, for instance, had reported its model extracted credentials from a real company and accessed a production database during its evaluation.
Further criticisms include the report's lack of concrete specifics such as dates, named owners for corrective measures, or verifiable criteria for improvements. Irregular's claims about monitoring tools being poorly suited for evaluation logs, while simultaneously identifying a significant expansion of manual review as a principal safeguard, are seen as contradictory. The company reiterated that there are "no active issues today" but also stated its audit remains underway.
The incidents and Irregular's response highlight broader scrutiny within the cybersecurity community regarding how AI companies and their contractors report containment failures during model evaluations. Security professionals argue that current disclosures fall short of industry standards. In contrast, the U.S. AI Security Institute's recent technical report on its own evaluation runs provided specific details, including model names, incident counts, timestamps, and a commitment to independent review, and confirmed notification of affected parties. Irregular's post did not offer comparable disclosures or confirm notification of affected third parties.
It remains unclear whether law enforcement or regulatory agencies have initiated investigations into these incidents, or if any affected third parties, who did not consent to being targeted and in some cases did not detect the intrusions themselves, are considering legal action. Neither Irregular nor its customers have publicly stated whether regulators were notified. Irregular has announced plans to publish an open white paper on best practices for evaluation security, including standards for internet access during pre-deployment testing, but has not provided a publication date.






