← All posts / Meta

Stealing Reasoning Traces: Researchers Decrypt the Encrypted Chain-of-Thought of Anthropic, OpenAI, and Google Models

A 116-page preprint shows encrypted reasoning blocks from frontier LLM APIs are interchangeable across sessions, users, and models — enabling a 'decryption jailbreak' that extracted 315,320 hidden traces, 367 PII artifacts, and 182 credentials from public repos.

Stealing Reasoning Traces: Researchers Decrypt the Encrypted Chain-of-Thought of Anthropic, OpenAI, and Google Models

Every major reasoning-model API — OpenAI’s, Anthropic’s, Google’s — hands clients an opaque Base64 blob alongside each response. It is the model’s hidden chain-of-thought, encrypted so the client can carry conversation state between turns without ever reading the model’s internal reasoning. The design intent is reasonable: providers protect their intellectual property, keep unfiltered intermediate thinking away from users, and still support stateless APIs, zero-data-retention modes, and tool loops.

On August 10, 2026, a research team published a 116-page preprint demonstrating that this design is broken. The paper, Stealing Reasoning Traces from Proprietary LLM APIs (arXiv:2608.09867), was written by Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, and Maksym Andriushchenko of the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems. Its core finding is architectural rather than a coding bug: encrypted reasoning blocks are fully compatible and interchangeable across different sessions, users, and models within a provider’s ecosystem.

How the decryption jailbreak works

The exploit is elegantly indirect. You do not attack the frontier model’s safety training at all. Instead, you take an encrypted reasoning block produced by a strong, well-guarded model and inject it into a conversation with a weaker sibling model from the same provider. The provider transparently decrypts the block on the way in — that is how the protocol works. The weaker model then sees the frontier model’s plaintext reasoning as legitimate context and can be prompted to transcribe it verbatim.

The strong model’s guardrails are never engaged. A smaller, cheaper, less-defended model in the same family does the reading. Per the paper’s own framing, Haiku 4.5 can read Opus 4.8’s thoughts. The researchers report the same trick working across Anthropic, OpenAI, and Google.

The cross-account portability also implies something uncomfortable about key management: providers are almost certainly encrypting every client’s reasoning under a single global key rather than per-account or per-session keys. If keys were scoped to an account, a block from account A would simply fail to decrypt under account B. That distinction matters most for customers relying on zero-data-retention guarantees — ZDR means the provider is not storing your conversation, but under a global key, your reasoning state is not cryptographically isolated from every other customer’s.

Four attack vectors

The paper enumerates four consequences of the flaw:

  1. Anti-distillation bypass. Extract a proprietary model’s raw reasoning traces at scale — demonstrated across all three providers. Hidden chain-of-thought was one of the main defenses against competitors harvesting frontier reasoning to train rival models; the paper says that defense had a hole in it.

  2. Private data extraction at scale. Developers constantly share full agent session logs publicly — in bug reports, gists, blog posts, and committed debug output — treating the encrypted blobs as unreadable noise. They were never unreadable. The researchers scraped public repositories, decoded 315,320 reasoning blocks, and recovered 367 PII artifacts and 182 credentials. Worse, the reasoning block can contain material the visible transcript does not: a model that reads a .env file, reasons about the key it found, and writes an answer that never mentions the key has still put that key into its chain-of-thought.

  3. Hazardous information disclosure. The hidden reasoning can contain unsafe content even in cases where the model’s final, visible output correctly refuses a malicious request.

  4. Invisible prompt injection. Attackers can embed malicious payloads entirely inside encrypted blocks, where they are invisible to human review and most log scanners — the blob is expected to be unreadable noise. The payload then travels through public agentic rollouts looking exactly like normal API bookkeeping.

The disclosure story is its own finding

This flaw was reported before, and dismissed. On May 29, 2026, Johns Hopkins cryptographer Matthew Green published a write-up describing a weekend spent probing these blobs. His findings anticipated much of the paper: blocks could be replayed within a session, across sessions, and across entirely separate accounts; for OpenAI they replayed across different models too. Green even demonstrated that the blocks are semantically active rather than inert — in one case, a social security number reasoned about in one session reappeared, unprompted, in a different session on a different account after the block was replayed.

Green reported both results through the providers’ bug bounty programs. Per his account, OpenAI called the report unreproducible; Anthropic said it saw no security implications in side channels or replays. Ten weeks later, the preprint demonstrated credential extraction at scale, cross-model reasoning theft, and invisible prompt injection built on the same underlying flaw.

Green’s structural analysis holds up as useful context: Anthropic’s block includes a field labeled “signature” that contained no actual signature; the 12-byte IVs suggest GCM or ChaCha and are probably too short; OpenAI’s format appeared to be based on the Fernet token standard. The decisive asymmetry: tampering with any ciphertext byte produces a clean API rejection, while replaying an unmodified block produces no error at all.

There is also a side channel that no cryptography can patch. If an attacker induces secret-dependent reasoning — cheap computation when a hidden bit is 0, expensive when it is 1 — the difference shows up in reasoning token counts, block length, and wall-clock response time. Green demonstrated bit-by-bit extraction of a byte this way. The leak is in how much the model thinks, not in what it wrote down.

What providers should fix — and what you should do today

The paper’s proposed mitigations follow directly from the failure mode, and none are exotic: bind blocks to a session identifier (kills cross-session replay), derive keys per account (kills cross-account replay and fixes ZDR isolation), bind blocks to the issuing model (kills the cross-model decryption jailbreak), rotate keys, and enforce nonce or sequence ordering. Scoping a token to the session and identity that produced it is standard practice; its absence is the gap.

For teams building on these APIs, three durable takeaways:

  • Encrypted does not mean isolated. Opacity to you is not confidentiality from everyone. A blob you cannot read may still be readable by anyone who can replay it into a cooperative model.
  • Agent session logs are now sensitive artifacts. Search your repos, gists, and issues for pasted agent sessions — look for long Base64 blobs in thinking, reasoning, or signature fields — rotate anything the session touched, and strip reasoning blocks before publishing anything.
  • Hidden reasoning is a weaker IP moat than it looked. If chain-of-thought concealment is part of your competitive defense, this paper narrows it considerably.

As of publication, the labs have not publicly detailed their remediation status — check their security advisories. But the deeper lesson predates any patch: when three frontier labs independently ship the same cryptographic shortcut, the industry’s security review process for novel protocol designs is not keeping pace with how fast these systems ship.