OpenAI ties a distillation campaign to Moonshot AI
OpenAI published Disrupting a coordinated model-distillation campaign on 30 September 2026, at 18:30 SGT by its news feed. The post says OpenAI found and disrupted a campaign to extract protected reasoning from its models.
OpenAI says it attributes "a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi." It adds that it is unclear whether all the operators came from a single actor.
OpenAI defines adversarial distillation as "the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model."
Timeline and scale from OpenAI's post
OpenAI gives this timeline:
- The activity began on 1 July 2026 at low volume.
- On 24 and 25 July, OpenAI saw spikes of 16,000 requests using an extraction pattern, from over 4,000 users.
- A footnote says these figures describe "attempted, not necessarily successful, extractions."
- OpenAI linked related prompt patterns across more than 15,000 users and says it fully disrupted that cluster by 28 July.
OpenAI says the operators did not break its encryption, compromise a database or reach stored user conversations. It says they manipulated model interactions so that protected reasoning appeared in a form the requester could see.
Encrypted reasoning was the target
OpenAI says operators copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe it. The post does not publish prompts, and this article does not describe the method further.
OpenAI says independent researchers reported related cross-model and conversation-compaction flaws through responsible disclosure. It says it confirmed "that the attack paths they identified were real." The post links the arXiv paper Stealing Reasoning Traces from Proprietary LLM APIs, submitted on 10 August 2026.
The paper's abstract says encrypted reasoning blocks are interchangeable across sessions, users and models within one provider. It says the authors demonstrated reasoning extraction across Anthropic, OpenAI and Google.
OpenAI lists these responses:
- It banned or restricted fraudulent accounts and tightened signup and infrastructure controls.
- It closed a pathway that let someone holding another user's encrypted reasoning replay it and recover the contents.
- It added checks to detect and hold streamed output that might expose reasoning.
- It worked with third-party services when related activity moved through them.
- It shared findings through the Frontier Model Forum and government information-sharing channels.
OpenAI's attribution has no outside confirmation
The attribution rests on OpenAI's own telemetry. The post names no government or industry body that has confirmed it.
When The Frontier checked Kimi's official X account at 07:32 SGT on 3 October, it showed no posts since 30 September. CNBC reported on 30 September US time that Moonshot did not immediately respond to its requests for comment.
OpenAI says the work is not finished. It says partner-hosted deployments need the same protections as its own services. It also says "Systems that support portable or replayable reasoning artifacts may face related risks."
Treat encrypted reasoning blocks as sensitive data
Teams that store or pass back encrypted reasoning from any provider should treat those blocks as sensitive. The arXiv abstract says developers often share session logs publicly without knowing what the encrypted blocks contain.
- Keep encrypted reasoning out of public logs, bug reports and shared transcripts.
- Ask cloud partners that host frontier models whether they have applied the provider's new reasoning protections.
- Watch for unusual request patterns that repeat reasoning blocks across accounts.
Anthropic made its own distillation claims against Moonshot and other labs in September. See Anthropic's Sept threat report is not one verified ledger. OpenAI warns that distilled models may lose the original's safeguards. See Anthropic says Z.ai GLM-5.3 brings Mythos-class cyber skills with weak open-weight safeguards.
