← All posts / Research

Zero-Click Grok Hack: 'Cryptographic Context Injection' Steals Chat Histories and xAI Still Hasn't Patched It

Adversa AI disclosed a zero-click attack that hides AES-256-GCM-encrypted instructions in ordinary web pages; when Grok summarizes them, its own Python sandbox decrypts the payload and exfiltrates the user's name, location, and full chat history — reported to xAI on June 3, still unfixed on August 19.

Zero-Click Grok Hack: 'Cryptographic Context Injection' Steals Chat Histories and xAI Still Hasn't Patched It

For years, the standard answer to prompt injection has been “filter the input.” On August 20, 2026, security firm Adversa AI published research that explains, in painstaking detail, why that answer is running out of road. The technique — dubbed Cryptographic Context Injection — smuggles malicious instructions into an AI agent as AES-256-GCM ciphertext that no content classifier can read, then lets the agent’s own Python sandbox decrypt the payload. From the model’s point of view, the attack doesn’t arrive from the outside at all. It emerges as the output of code the model just ran, and the model treats it as trusted internal state.

The result, demonstrated against xAI’s Grok web chat at grok.com running Grok 4.5 Fast, is a genuinely zero-click data theft: a user asks Grok to summarize an ordinary-looking web page, and the agent quietly transmits the user’s name, approximate location, subscription tier, and the complete set of prompts in the ongoing conversation to an attacker-controlled server. No confirmation dialog. No visible warning. And as of August 19 — more than ten weeks after Adversa first reported the flaw to xAI — no patch.

How the attack works

The mechanics deserve careful attention, because they mark a genuine evolution in the prompt-injection arms race.

Step 1: The payload ships as ciphertext. The attacker hosts a web page containing an encrypted JSON object, the key material needed to decrypt it, and an innocuous instruction telling the agent to decrypt the blob using its Python runtime. Everything a guardrail scanner would need to catch the attack is right there on the page — but recovering the plaintext requires actually executing PBKDF2 key derivation and AES-256-GCM decryption, which no content classifier does at inspection time. Input filters classify text; they do not run it.

Step 2: The sandbox launders the payload into trusted context. This is the heart of the technique, and the reason it differs from prior cipher-based evasion attacks. Earlier research — CipherChat, CodeChameleon — used weak, reversible encodings like substitution ciphers, XOR, or base64, which models can decode natively “in-weights” from their training data. Strong encryption cannot be decoded that way. It forces recovery through the code execution runtime. Once Grok’s sandbox runs the decryption, the attacker’s instructions exist not as an untrusted fetched string, but as the return value of the model’s own code. Adversa calls this the “trust laundering channel,” and the closest classical analogy is SQL injection: a system failing to distinguish its own trusted operations from attacker-supplied data flowing through the same channel.

Step 3: A disguised “decryption key” carries the loot. The decrypted instructions direct the agent to build an additional “decryption key” that is not key material at all. Its value is a template string that interpolates the user’s private session context — name, coarse location, subscription tier, and full conversation prompts. The agent is then told to open a URL “to fetch additional context,” and it autonomously invokes its privileged navigation tool with that URL, the exfiltrated data riding along in the query parameters.

In the proof-of-concept, the chain completed with no user confirmation and no visible warning — triggered by nothing more suspicious than “summarize this page.”

The numbers behind the disclosure

Adversa tested the attack roughly 20 times since June with a 40% success rate — and notably, the failures came from Grok struggling with the decryption itself, not from any prompt or response being flagged by safety systems. The company reproduced the attack as recently as August 19, 2026, one day before publication.

The disclosure timeline is the story’s most uncomfortable detail:

  • June 3, 2026 — Adversa reports the attack to xAI directly and through its HackerOne bug bounty program.
  • xAI acknowledges the report but provides no specifics and no mitigation timeline.
  • August 4 and August 10 — Adversa makes two more attempts to coordinate. No response.
  • August 20 — Adversa discloses publicly, withholding operational payloads to prevent exploitation.
  • As of publication: no patch, no CVE identifier, no user-facing workaround, and no xAI advisory.

Gemini, too — with a twist

The same writeup demonstrates a second variant against Google’s Gemini public chat interface in Deep Thinking mode (identified via extracted system instructions as Gemini 3 Flash, Web, paid tier). Because Gemini’s Python sandbox has no access to external websites, the exfiltration scenario doesn’t transfer — but the policy bypass does.

In that version, a single prompt asks Gemini to decrypt a ciphertext whose plaintext is a convincingly formatted fake Python traceback. The “traceback” is the payload: it contains a fabricated, API-styled callback that appears to deactivate safety policy and a first-person reasoning prefix that pre-commits the model to producing the restricted content, framed as something it will encrypt “for safety.” The model generates the prohibited answer in plaintext, encrypts it inside the sandbox, and the output filter — seeing only ciphertext — waves it through. The technique produced a multi-paragraph description of how to build an incendiary weapon, content Gemini’s filters normally suppress.

Google was never notified, because jailbreaks fall outside its vulnerability disclosure program’s scope. Adversa reports the success rate against Gemini has dropped sharply by August, though it can’t attribute whether that reflects filter updates or model changes. Interestingly, the Gemini chain isn’t entirely new: lead researcher Rony Utevsky published substantially the same demonstration on his personal site on March 11, 2026, under the name “Cryptographic Payload Injection,” reporting five out of five successful reproductions — and cross-model results in which GPT-5 failed to parse the decryption instructions while Claude Sonnet 4.5 decrypted the payload, then correctly flagged it as prompt injection.

Why this matters beyond Grok

The uncomfortable lesson is architectural, not incidental. The xAI framework lets instructions parsed from an untrusted external page drive a privileged, internet-connected tool; it resolves private session metadata into the inputs of that outbound tool; and it enforces no effective egress boundary, no consent gate, and no provenance separation on the path. That combination is exactly what Simon Willison’s “lethal trifecta” warned about: agents with access to private data, exposure to untrusted content, and the ability to take outward-facing actions.

And chatbots are the least dangerous case. As Adversa notes, every precondition is stronger for coding agents, platform-operations agents, and financial agents — where code execution isn’t an exception path but the product, outbound network calls are routine, and the credentials in reach are worth far more than session metadata.

The defense guidance is similarly harness-centric, not model-centric:

  1. Quarantine untrusted content in a context with no tools and no credentials; return only structured data to privileged contexts.
  2. Gate irreversible and outbound actions — confirm new network destinations, pushes, merges, and publishes, showing fully resolved arguments rather than templates.
  3. Capture per-session tool traces with resolved arguments; without them there is neither detection nor forensics.
  4. Alert on sequences, not payloads — untrusted content enters, code executes, then the agent contacts a host outside its dependency graph. An opaque blob paired with decryption instructions is a review signal, not a blocking filter.
  5. Make context provenance a procurement requirement — ask vendors whether tool output is separated from the instruction channel.

The bottom line

Cryptographic Context Injection is not the first prompt-injection attack, and it won’t be the last. What it proves is that the attack surface has quietly expanded from “model inputs” to everything an LLM treats as its own context — tool outputs, runtime results, intermediate state. Guardrails that only inspect text on the way in are structurally blind to instructions that materialize inside the runtime. Until xAI ships a fix, the practical advice for Grok users is blunt: think twice before asking the agent to browse and summarize pages you don’t control. The page might be asking Grok to summarize you right back.