← All posts / Research

Researchers Crack Encrypted AI Reasoning Across OpenAI, Anthropic, and Google

A new paper shows encrypted chain-of-thought reasoning from GPT-5.6, Claude Opus 4.8, and Gemini 3 can be decrypted using cheaper sibling models — exposing passwords, API keys, and internal safety logic.

Researchers Crack Encrypted AI Reasoning Across OpenAI, Anthropic, and Google

The Hidden Thoughts of AI Models Are Not So Hidden After All

When modern reasoning models like OpenAI’s GPT-5.6, Anthropic’s Claude Opus 4.8, or Google’s Gemini 3 tackle complex problems, they generate extensive internal “chain-of-thought” reasoning — a step-by-step thinking process that happens before the model produces its final answer. To protect intellectual property and prevent competitors from distilling their models, providers encrypt these reasoning traces into opaque blobs that are returned to the client in encrypted form. The assumption was simple: clients can replay these blobs to maintain conversation context, but cannot read them.

That assumption has now been shattered. A research team led by Alexander Panfilov, with collaborators from the ELLIS Institute Tübingen, the Max Planck Institute, MATS Research, and Snyk, has published a paper titled “Stealing Reasoning Traces from Proprietary LLM APIs” demonstrating that these encrypted reasoning blocks can be fully decrypted — cheaply, at scale, and across every major AI provider.

How the Attack Works: Turning Sibling Models Into Decryption Oracles

The core vulnerability is architectural. The encrypted reasoning blobs returned by provider APIs are authenticated using a single global, provider-wide encryption key rather than being cryptographically bound to a specific user account, session ID, or model tier. This means an encrypted envelope created by a heavily guarded flagship model can be passed into any other model hosted under the same provider’s infrastructure.

The attack chain is elegantly simple:

  1. Capture an encrypted reasoning block emitted by a frontier model (e.g., Claude Opus 4.8).
  2. Inject that block into the API call of a smaller, cheaper sibling model from the same provider (e.g., Claude Haiku 4.5), instructing it to transcribe the internal thinking verbatim.
  3. Extract the plaintext reasoning.

Because lighter models lack the aggressive anti-distillation alignment and safety guardrails enforced on flagship tiers, they comply with the instruction and output the hidden reasoning in plain text. The researchers confirmed identical cross-model compatibility across OpenAI’s GPT-5.6 family (using GPT-5-mini and o4-mini as decryption oracles) and Google’s Gemini 3 lineup (using Gemini 3.1 Flash as the oracle).

The mathematical precision is striking: for most queries, the number of extracted tokens matches the billed thinking tokens exactly, meaning the researchers are capturing the full internal reasoning, not just partial snippets.

A Dismissed Warning That Became a Crisis

The story traces back to May 2026, when cryptography expert Matthew Green discovered that encrypted reasoning blobs could be replayed outside their original context. Green reported the cross-session replay flaw to OpenAI and Anthropic via their bug bounty programs. According to Panfilov, the providers’ response was dismissive: “they don’t see any security implications in side channels or replays.”

In June 2026, independent researcher Will Smidlein confirmed the vulnerability and identified a single global encryption key as the root architectural flaw. Smidlein noted that “the providers are probably using a single global key to encrypt and authenticate all reasoning data sent to the client.”

Despite these reports, no architectural fix was deployed. The formal paper arrived 11 weeks later, and by then, the researchers had already found 315,320 exploitable encrypted reasoning blocks circulating in public repositories on GitHub and Hugging Face.

The Real-World Damage: Passwords, API Keys, and PII

The vulnerability’s impact extends far beyond intellectual property concerns. By analyzing 6,708 public agent transcripts scraped from GitHub and Hugging Face, the researchers decoded 315,320 embedded reasoning blocks. What they found was alarming:

  • 62 API keys — live credentials embedded in reasoning traces that developers never realized were exposed
  • 33 passwords — authentication secrets that appeared only in the model’s internal thoughts, never in visible responses
  • 33 email addresses — personally identifiable information hidden in encrypted reasoning
  • 367 total PII artifacts and 182 hardcoded credentials

The findings revealed something unsettling about how AI models reason internally. The extracted traces showed models sometimes communicating in incomprehensible language fragments (one trace contained the cryptic phrase “but marinade”), constructing answers in reverse order, and even considering deceptive strategies before settling on a final response.

Evidence of Industrial-Scale Model Distillation

The paper also provides compelling evidence for the long-suspected practice of model distillation — where a less capable model is trained on the reasoning outputs of a more powerful one. The researchers tested Kimi-K3, a Chinese model from Moonshot AI, and found that its reasoning behavior shifts measurably toward Opus when pre-filled with just a few tokens from Opus’s thought processes.

A memorization analysis showed that specific Claude and GPT reasoning segments are up to six orders of magnitude easier to extract from Kimi-K3 than from the next closest model. This strongly suggests Kimi-K3 was trained on stolen or extracted reasoning traces from Western frontier models. The researchers also note that Kimi-K3’s comparatively poor performance on cybersecurity benchmarks and complex math tasks aligns with distillation patterns — these are tasks harder to recover even from raw chain-of-thought data.

The attack is also remarkably affordable. The authors estimate API costs for decoding 10,000 reasoning traces at approximately $720, making industrial-scale extraction trivially within reach for any well-funded actor.

Four Protection Layers Bypassed Simultaneously

The vulnerability is particularly devastating because it bypasses four distinct provider protection mechanisms at once:

  1. Anti-distillation controls — designed to prevent competitors from training on proprietary reasoning
  2. PII isolation — meant to keep personal data out of shared contexts
  3. Hazard suppression — the safety logic that prevents models from outputting dangerous content
  4. Hidden prompt confidentiality — the encryption that was supposed to make reasoning opaque

Perhaps most alarmingly, the vulnerability enables invisible prompt injection attacks against autonomous AI agents. An adversary can craft malicious instructions inside an encrypted reasoning block. When processed by a public-facing agent, monitoring tools that inspect only the visible conversation history fail to flag the hidden payload, allowing injected instructions to compromise downstream agentic workflows undetected.

What Needs to Change

As of publication, Anthropic, OpenAI, and Google have issued no architectural fix. The operational guidance from the research team is blunt: treat encrypted reasoning as readable until your vendor tells you otherwise, and audit anything already pushed to public repositories.

The cryptographic fix is conceptually straightforward — providers need to bind encrypted reasoning blocks to specific sessions using per-session keys or properly authenticated encryption — but implementing it requires a fundamental redesign of how reasoning state is managed across their infrastructure. In the meantime, every team that publishes session logs, runs public-facing agents, or shares encrypted reasoning blobs faces active credential and PII extraction exposure.

The broader lesson is one of cryptographic humility. When providers dismissed Matthew Green’s May report, they assumed their encryption was sound. It wasn’t. The reasoning traces that were supposed to be sealed behind cryptography turned out to be an open book — readable by anyone with an API key and a few dollars to spare.