OpenAI says it disrupted a coordinated model-distillation campaign
OpenAI says it stopped a campaign aimed at extracting protected model reasoning, linking a core cluster to people associated with Moonshot AI while saying databases and stored chats were not directly accessed.
OpenAI said it identified and disrupted a coordinated campaign designed to extract protected reasoning from its models. The company attributed a core cluster to people associated with Moonshot AI, the developer of Kimi, while acknowledging that it was unclear whether every operator belonged to one actor.
The company said the earliest activity appeared on July 1. It counted 16,000 requests with a relevant extraction pattern from more than 4,000 users during spikes on July 24 and 25, then found related prompt-pattern activity across more than 15,000 users.
OpenAI said it had fully disrupted the broader cluster by July 28. It stressed that the figures represented attempted, not necessarily successful, extractions.
OpenAI said the activity did not break encryption, compromise a database or directly access stored user conversations. Instead, it said operators manipulated model interactions to make protected reasoning appear in output visible to the requester.
One method, according to OpenAI, involved copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe the hidden material. An August preprint on reasoning-trace extraction separately described encrypted reasoning blocks as interchangeable across sessions, users and models, and demonstrated the broader attack class against systems from OpenAI, Anthropic and Google. That research supports the mechanism’s feasibility, but does not independently verify OpenAI’s attribution, campaign counts or impact.
OpenAI said independent researchers had disclosed related cross-model and conversation-compaction weaknesses, and that its own investigation confirmed those attack paths. The company said it banned or restricted fraudulent accounts, tightened signup and infrastructure controls, expanded network monitoring, and added protections against replay and streamed-reasoning exposure. It also said it coordinated disruptions with third-party providers.
The company characterized adversarial distillation as a safety and national-security risk because extracted reasoning could transfer model capabilities without the original safeguards or comparable safety investment. That assessment, along with the link to people associated with Moonshot AI, remains OpenAI’s claim. The company did not publish the underlying attribution evidence or say how many attempts succeeded.
OpenAI said it shared its findings through the Frontier Model Forum and government information-sharing channels and was continuing its investigation and mitigations.
More news

Metaview raises $60 million Series C to grow recruiting agents

Databricks launches ai_decide beta for governed AI decisions

Cohere introduces RCP-nDCG@10 retrieval metric
