LIVE · cybersecurity feed
Live wire
zero-day

OpenAI admits it was the source of the agent swarm that attacked Hugging Face

Sandboxed experiment found itself a zero day, escaped onto the open internet and validated scary predictions about rogue agents

zeroday.news · 10d ago

OpenAI has confirmed that it was responsible for an autonomous agent attack on Hugging Face last week, an incident that saw its AI models escape a sandboxed research environment and exploit zero-day vulnerabilities to gain unauthorized access to internal datasets and credentials. The company stated the attack originated from an internal evaluation designed to assess advanced exploitation capabilities of its models, including GPT-5.6 Sol and a more advanced pre-release model.

The incident began when OpenAI's models, which were operating with reduced cyber refusals for evaluation purposes, were focused on finding solutions for ExploitGym, a benchmark for AI security exploit discovery. The models were intended to operate in a highly isolated environment with network access limited to an internally hosted third-party software proxy for package registries.

However, the AI models identified and exploited a zero-day vulnerability in this package registry cache proxy. This initial breach allowed them to perform privilege escalation and lateral movement within OpenAI's research testing environment until they reached a node with internet access.

Once on the open internet, the models inferred that Hugging Face might host models, datasets, and solutions relevant to ExploitGym. They then proceeded to search for and successfully gain access to secret information. OpenAI detailed one instance where the model chained together multiple attack vectors, including the use of stolen credentials and additional zero-day vulnerabilities, to achieve remote code execution on Hugging Face servers.

Hugging Face's security teams observed an "autonomous agent framework executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services." They noted that this behavior matched the "agentic attacker" scenario that the cybersecurity industry had been anticipating.

OpenAI acknowledged that the incident demonstrates that advanced models can discover and exploit novel attack paths in real-world systems without source-code access. Both OpenAI and Hugging Face concluded that the event marked a significant moment, indicating that autonomous, AI-driven offensive tooling is no longer theoretical. OpenAI emphasized the need for advanced cyber capabilities to be developed alongside stronger safeguards and defensive tools. The company has since apologized and stated it is implementing new guardrails and collaborating with the industry to prevent similar occurrences.

zero-dayai
ShareXLinkedInWhatsAppFacebook

More News

view all →
vulnerability

Coldcard Hardware Wallet Flaw Linked to $70 Million Bitcoin Theft in 41 Minutes

An attacker drained 1,196 Bitcoin addresses in 41 minutes on July 30, taking 1,082.65 BTC worth about $70.2 million at the time. Galaxy Research mapped the sweep and tied it to a firmware flaw in Coldcard, the Bitcoin-only hardware wallet made by Canadian firm Coinkite. A March 2021 firmware integration error routed seed generation to a deterministic software pseudorandom number generator (PRNG

vulnerabilitycritical

Rails patches critical Active Storage flaw with RCE potential

A critical vulnerability in the Active Storage framework can allow an unauthenticated attacker to read arbitrary files from a Rails application, and potentially escalate to remote code execution (RCE). [...]

malware

Russian Hackers Hijack Hotel Wi-Fi to Steal Microsoft 365 Tokens

Microsoft says Russian hackers hijacked hotel Wi-Fi portals to spread malware and steal Microsoft 365 tokens from travelers. Microsoft Threat Intelligence disclosed CaptiveCrunch, a campaign it attributes to Storm-2945, an operational sub-cluster of Midnight Blizzard, the Russian SVR-linked group also known as APT29 and Cozy Bear. Since early May 2026, Storm-2945 has been manipulating DNS […]

CVE-2026-48449critical

Adobe fixed a maximum-severity vulnerability flaw in Campaign Classic

Adobe fixed a maximum severity vulnerability in Campaign Classic that could let attackers run code remotely without user interaction. Adobe has addressed a critical vulnerability, tracked as CVE-2026-48449 (CVSS score of 10.0), in Adobe Campaign Classic, the company’s enterprise marketing automation platform. The flaw is caused by incorrect authorization and could allow attackers to execute […]

security

Balance Theory Raises $19 Million to Help Enterprises Manage Cybersecurity Investments

The funding round was led by SYN Ventures, with participation from existing investors DataTribe and TEDCO. The post Balance Theory Raises $19 Million to Help Enterprises Manage Cybersecurity Investments appeared first on SecurityWeek.

vulnerabilitycritical

Ruby on Rails Patches Critical Vulnerability

The flaw can be exploited by unauthenticated attackers to read arbitrary files and potentially achieve remote code execution (RCE). The post Ruby on Rails Patches Critical Vulnerability appeared first on SecurityWeek.