OpenAI has disclosed that it disrupted a sophisticated campaign aimed at extracting reasoning capabilities from its AI models, attributing a significant portion of the activity to individuals associated with the Chinese company Moonshot AI. The company described the attack method as "novel," involving an encryption bypass that did not compromise its systems directly but rather manipulated model interactions.
The activity was first detected on July 1, 2026, with low-level probing that escalated significantly by July 24 and 25. During this peak period, OpenAI observed approximately 16,000 prompts originating from 4,000 users exhibiting a consistent extraction pattern. By July 28, the number of suspicious users had grown to 15,000, at which point OpenAI stated it had fully disrupted the operation.
The attackers' technique involved copying encrypted reasoning data from one conversation and then, in a separate conversation, prompting the AI model to decrypt and transcribe this content in plain text. OpenAI clarified that this method did not involve breaking their encryption, compromising databases, or gaining unauthorized access to stored user conversations. Instead, it exploited a vulnerability in how the model handled encrypted data across different conversational contexts.
OpenAI stated that while it is unclear if all observed activity is linked, a core cluster of the attacks was attributed to individuals working on behalf of Moonshot AI. Moonshot AI is a China-based rival AI company whose Kimi model is among several Chinese open-source AI offerings. OpenAI's public statement did not provide specific technical evidence to support this attribution.
The company confirmed that outside researchers had independently reported a similar vulnerability to them in August. OpenAI also indicated that this vulnerability is not exclusive to its models and has shared information about the incident with organizations like the Frontier Model Forum.
In response to the incident, OpenAI implemented several security enhancements. These include banning the accounts involved in the attack, improving signup and infrastructure controls, expanding network monitoring capabilities, and fixing the specific bug that allowed encrypted data from one conversation to be decrypted in another.






