← All posts / Meta

16,000 Requests in 48 Hours: OpenAI Disrupts Moonshot-Linked Campaign to Steal Its Models' Hidden Reasoning

OpenAI's September 30 disruption report details a coordinated adversarial-distillation campaign that peaked at 16,000 extraction requests from 4,000+ users in two days, attributing a core cluster to individuals associated with Moonshot AI — and exposes the encrypted reasoning-trace architecture every frontier lab now has to defend.

16,000 Requests in 48 Hours: OpenAI Disrupts Moonshot-Linked Campaign to Steal Its Models' Hidden Reasoning

On September 30, 2026, OpenAI published a security post titled “Disrupting a coordinated model-distillation campaign” — and in doing so pulled back the curtain on one of the most consequential infrastructure disputes of the current AI era: the fight over who owns a model’s hidden chain of thought. The company said it identified and disrupted a coordinated campaign designed to extract protected reasoning from its models, and it attributed a core cluster of the activity to individuals associated with Moonshot AI, the Beijing-based developer of the Kimi model family. The earliest observed activity dates to the first week of July 2026.

What actually happened

OpenAI described the activity as consistent with adversarial distillation — the systematic, unauthorized use of one model’s outputs or reasoning to train, reproduce, or improve another model. The target was not user data in the conventional sense. The operators did not break OpenAI’s encryption, compromise a database, or gain access to stored user conversations. Instead, they manipulated model interactions so that protected reasoning — the model’s internal record of working through a task — could be reproduced in forms visible to the requester, at coordinated scale, in violation of the company’s terms of service.

The timeline OpenAI disclosed is precise. Activity began on July 1, 2026, initially at low volume. On July 24 and 25, it spiked: 16,000 requests using a relevant extraction pattern, originating from more than 4,000 users in just two days. A footnote in the post is careful to note these figures describe attempted, not necessarily successful, extractions. Follow-up investigation identified related prompt-pattern activity across a cluster of more than 15,000 users, which OpenAI says it had fully disrupted by July 28, 2026. The activity evolved over the course of the campaign — a detail the company treats as evidence that adversarial distillation is a broader security challenge requiring layered, adaptive defenses, not a one-off incident.

The decryption trick at the center

The most technically interesting detail is how the extraction worked. Frontier providers conceal step-by-step reasoning to protect intellectual property, returning reasoning traces to the client as blocks of encrypted text that the client passes back with each subsequent request. One technique OpenAI observed involved copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe the hidden content.

OpenAI is careful to stress that this manipulation is not a vulnerability unique to its models. Independent security researchers — in a paper submitted to arXiv on August 10, 2026 by Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, and Maksym Andriushchenko — had already mapped the architectural weakness: encrypted reasoning blocks are fully compatible and interchangeable across different sessions, users, and models within a provider’s ecosystem. That interchangeability enables a “scalable decryption jailbreak” in which an encrypted trace from a capable model is injected into a weaker, less safeguarded model from the same provider, forcing it to decode and output the trace verbatim in plaintext — without ever jailbreaking the more capable model directly. The researchers demonstrate the attack class across Anthropic, OpenAI, and Google, and report four distinct vectors: circumventing anti-distillation mechanisms, large-scale private data extraction (they recovered 367 PII artifacts and 182 credentials by decoding 315,320 reasoning blocks scraped from public repositories), inadvertent revelation of hazardous information hidden in reasoning even when the visible output safely refuses, and invisible prompt injections embedded entirely within encrypted blocks. OpenAI says the responsible-disclosure process confirmed the attack paths were real and accelerated its mitigations.

Attribution, carefully hedged

OpenAI’s attribution language is deliberately narrow. It is “unclear whether all operators observed during the relevant period originated from a single actor,” and the company attributes “a core cluster” of the activity to individuals associated with Moonshot AI — not to Moonshot AI the corporation itself. Moonshot did not immediately respond to CNBC’s requests for comment.

The context, however, is hard to ignore. On September 8, 2026, the NSA, CISA, and FBI released joint advisory AA26-251A stating that Moonshot AI has conducted a widespread distillation campaign against U.S. frontier AI companies since at least mid-2025 — allegedly extracting significant Claude Fable 5 data to train Kimi-K3 and GPT-4o data to train Kimi-K2, and using U.S. models to distill supervised fine-tuning, reinforcement learning, software engineering, and math capabilities. More broadly, the advisory claims that — likely with Chinese government awareness — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI have extracted billions of tokens across millions of exchanges from Claude, GPT, Gemini, and Grok variants since at least late 2024, routing requests through native APIs, remote cloud providers, third-party aggregators, and a gray market of API proxies known as “transfer stations.” And just weeks before OpenAI’s post, Anthropic publicly accused Chinese AI developers including Moonshot and Alibaba of secretly using Claude to train their own systems.

How OpenAI responded

The mitigation effort ran on three tracks. Account enforcement: banning or restricting fraudulent accounts, and strengthening signup and infrastructure controls to make coordinated account farming harder. Technical controls: closing the pathway that allowed someone who already possessed another user’s encrypted reasoning to replay it and recover its contents, and adding checks to detect and hold streamed output that might expose reasoning — protections strengthened across users, workspaces, organizations, and model families. Partner coordination: working with third-party providers to identify and disrupt accounts when related activity moved through their services, and sharing findings through the Frontier Model Forum and government information-sharing channels.

Two forward-looking warnings stand out. First, OpenAI notes that partner-hosted deployments need the same protections as first-party services — a pointed admission that the same model served through a cloud marketplace or reseller can become the weak link. Second, the company flags that tool-output attacks require protections examining more than ordinary visible text, and that systems supporting portable or replayable reasoning artifacts face related risks as a class.

Why this matters beyond one incident

Strip away the geopolitics and this is a story about a structural fault line in how frontier models are architected. Reasoning traces are simultaneously (a) the crown jewels of the model — the closest thing to its internal algorithm that a competitor can observe — and (b) a client-side artifact that must cross trust boundaries by design, because the API is stateless and the reasoning must accompany each request. That tension is what the decryption-replay attack exploits, and it cannot be patched away with a classifier. OpenAI itself expects adversarial distillation attempts to become more sophisticated as frontier models improve and as actors look for cheaper ways to mimic capabilities.

The safety framing deserves scrutiny too, because it is doing real work in OpenAI’s argument. Extracted reasoning could be used to train another model without preserving the safeguards applied to the original model’s user-facing outputs — meaning the stolen capability arrives uncoupled from its guardrails. Distillation at scale lets an actor acquire advanced capabilities without the accompanying investment in safety, a concern that sharpens as models gain dual-use capabilities. Whether one accepts the national-security framing or not (The Register, for one, greeted the post with an “irony alert” over OpenAI’s own training-data practices), the operational point stands: the reasoning trace is now contested infrastructure, and every provider that returns encrypted reasoning to clients is defending the same attack surface.

For the industry, the episode consolidates a new norm in under three months: joint US intelligence advisories naming foreign AI firms, victim companies publishing disruption reports with user counts and request volumes, findings shared through the Frontier Model Forum, and mitigation playbooks propagating across cloud partners. Frontier AI security has started to look like an intelligence discipline — with attribution caveats, disclosure timelines, and coordinated response. The 15,000-account campaign is disrupted. The attack class it exploited is not going anywhere.