OpenAI says it cut off a reasoning-theft campaign, but the same trick worked on Azure
OpenAI says it shut down an adversarial distillation campaign that pulled hidden reasoning out of its models, but independent researchers found the same method still worked on Microsoft Azure as late as 13 September, according to a study update published on 1 October.

OpenAI published a blog post on 1 October saying it detected and disrupted an "adversarial distillation" campaign, described as the systematic and unauthorised use of one model's outputs or reasoning to train or improve another. The company links a core cluster of the activity to people associated with China-based Moonshot AI, the developer of Kimi. Moonshot did not immediately respond to CNBC's request for comment.
OpenAI says the activity began at low volume on 1 July, then spiked on 24 and 25 July to 16,000 requests from more than 4,000 users, all following a typical extraction pattern. A wider look uncovered a network of more than 15,000 accounts with related patterns, which the company says it had fully shut down by 28 July. A footnote in the post says these were attempted extractions, not necessarily successful ones.
The mechanism matters more than the headcount. According to OpenAI, attackers took encrypted reasoning from one conversation and asked a model in a separate conversation to decrypt it and write it out.
Same models, different protections
Researchers Joachim Schaeffer and his team had already shown in the paper "Stealing Reasoning Traces from Proprietary LLM APIs", dated 10 August, how this works. AI providers send reasoning back to customers as encrypted packets, which customers pass along with follow-up requests. Because those packets are encrypted with shared keys, the researchers found, they can be moved between sessions, between users and even between different models from the same provider. That lets a weaker, cheaper model from the same family act as a "decryption oracle" that prints the stronger model's hidden thoughts word for word.
OpenAI credits the researchers by name and says their findings helped it roll out countermeasures faster. The company says it has banned fraudulent accounts, tightened sign-ups and closed the hole that let people reuse encrypted reasoning that did not belong to them. It also screens streamed outputs and holds them back if they might reveal reasoning, and says it shared what it learned through the Frontier Model Forum and government channels. OpenAI says its encryption was not broken, no database was compromised and there was no direct access to stored user conversations.
The researchers' update, published the same day, says the problem moved rather than disappeared. When the team tested again on 13 September, the attack was blocked on OpenAI's and Anthropic's own APIs. On Microsoft Azure it worked against every OpenAI model they tried, including GPT-6 Astra, and against Anthropic models up to Sonnet 5. A single attempt was enough to pull out verbatim reasoning. "Same models, but different protections depending on which platform serves them," Schaeffer wrote on X.
"We stole reasoning. Again." Joachim Schaeffer, researcher, in a post on X on 1 October
According to the researchers' timeline, OpenAI did not add safeguards to the Azure endpoint until 27 September. For Anthropic models, the reported extraction could no longer be reproduced on Azure starting 28 September. The researchers describe the fixes so far as piecemeal and superficial, with many relying on brittle matching of specific request patterns. GPT-6 Astra, they note, launched on third-party platforms without any of the protections in place.
Schaeffer argues that patches have to cover every type of attack and every cloud that hosts the models. Otherwise, attackers can simply pick the route with the weakest defences.
A second, simpler route
There is also a second method, demonstrated publicly by developer Can Bölük. The model gets a virtual notepad as a tool and is told to write its reasoning there, and the user can then read whatever it wrote. The researchers say this worked on every OpenAI model, as well as on Opus 4.8 and Sonnet 5. Only Opus 5, Fable 5 and Fable 5.1 did not reveal their reasoning. The output closely resembled what the decryption attack produced, and the researchers believe it would likely be just as useful for distillation.
OpenAI is not the only company reporting this. Anthropic's report "Detecting and countering misuse of AI: September 2026" described a single ten-day period in which Moonshot relayed almost 300,000 customer requests to Anthropic through a proxy network of 5,380 fraudulent accounts, according to Tom's Hardware. That report says Moonshot saved Claude's reasoning signatures and, in new sessions, got Claude to convert them back into full reasoning traces, which Anthropic calls cross-session replay attacks. In July, Moonshot denied that Kimi K3 was created from distillation.
The disclosures land in a crowded week for OpenAI. On 29 September the company unveiled "dots", an always-on assistant powered by GPT-6 Astra, less than 24 hours after saying it would scrap the launch of GPT-6.1 Astra over deceptive behaviour found in testing, the Guardian reported. OpenAI also apologised for one of its agents hacking an Australian government website.
For open-weight developers, the practical lesson is narrower than the policy debate. If a provider sells access to a model through a cloud endpoint, the protections on that endpoint, not the model's own API, decide whether hidden reasoning can be lifted.
Sources
5- 01OpenAI says it stopped a campaign to steal its models' reasoning, but the trick still worked on AzureEN
- 02OpenAI says actors linked to China-based Moonshot AI spearheaded a campaign to extract its models' hidden reasoningEN
- 03AI race heats up as OpenAI flags alleged model-copying campaignEN
- 04OpenAI announces 'dots' agent after scrapping launch of new AI model over safety concernsEN
- 05OpenAI unveils AI assistant 'dots' while safety worries delay new modelEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.