Claude Now Writes With an Invisible Watermark — Here's How the SynthID-Style Scheme Actually Works
Anthropic is embedding imperceptible watermarks in all Claude-generated text worldwide to comply with the EU AI Act, using a Google DeepMind-derived technique that bends randomness instead of words.
Every word Claude writes now carries a signature you will never see. In August 2026, Anthropic began embedding invisible, machine-readable watermarks in all text generated by Claude models launched on or after August 2 — and, in a decision with consequences far beyond Europe, it applied the marking worldwide, not just to EU users.
The move makes Anthropic one of the first major frontier labs to operationalize the EU AI Act’s transparency requirements for text, alongside roughly 190 signatories of the bloc’s Code of Practice on Transparency of AI-Generated Content. But the more interesting story is technical: how do you watermark language itself, without changing what the language says?
The Legal Trigger: Article 50(2)
As of August 2, 2026, AI providers serving the EU market must mark AI-generated content in a machine-readable format. Article 50(2) of the AI Act — with its transparency Code of Practice signed in July 2026 — obliges providers of systems that generate synthetic text, images, audio, or video to ensure their outputs are marked and detectable.
Anthropic’s obligation is technically narrower than Google’s, since Claude does not natively generate audio or video. But for text — the modality Claude dominates in coding and enterprise workflows — the company chose to go all-in: marking applies across the Claude Platform API, the Claude chat interface, Claude Code, Claude Cowork, and Claude Tag, on every supported surface, in every region.
Crucially, this includes output served through cloud partners — AWS, Google Cloud, and Microsoft Foundry. If a supported model is doing the writing, the watermark is in the words.
How the Watermark Works: Rigging the Dice, Not the Words
Anthropic’s own explanation, published August 14, is refreshingly concrete. Large language models generate text one token at a time, choosing among a list of candidate next-words. Often, several candidates are essentially interchangeable — given “The weather today was cold and…”, the next word could plausibly be “overcast” or “grey” without any change in meaning. Under normal operation, that choice is settled by a random number.
Watermarking changes only the source of the randomness.
Instead of an arbitrary random number generator picking between “overcast” and “grey”, the model uses a cryptographic key plus the preceding words to settle the choice. The words are still picked at random from the reader’s perspective — but someone holding the key can scan a passage and compute the probability that the sequence of choices is consistent with what Claude would have produced under that key. The pattern is statistical, not literal: the model isn’t always biased toward “overcast”; which word wins depends on what came before.
Anthropic is explicit about what the method is not. Nothing is added to the text. There are no hidden characters, no zero-width steganography, no invisible Unicode tricks. The watermark produces no extra tokens, so it costs nothing extra to serve and doesn’t slow the model down. And it carries no identifying information — nothing in the watermark or its key can be traced back to a specific person, organization, or conversation.
The lineage is notable. Claude’s watermark is a version of SynthID-Text, the technique Google DeepMind published in a Nature paper in 2024, which itself descends from Scott Aaronson’s 2022 watermarking proposal. The family shares one design principle: the watermark only changes the source of randomness, never the meaning of the words. DeepMind validated the approach in production, serving watermarked output to a slice of Gemini traffic and finding no statistically significant difference in user thumbs-up/thumbs-down ratings versus the unwatermarked model. Human raters comparing watermarked and unwatermarked answers side by side saw no difference in quality.
Where the Watermark Fades
Anthropic’s disclosure is unusually honest about the scheme’s limits — and this is where the story gets practical for anyone hoping to build detection into a workflow.
Short texts are hard. Detection works statistically; a paragraph has fewer word choices and thus less signal. Confidence that Claude was involved grows with length.
Factual passages are sparsely marked. Where only one word is correct — after “Isaac Newton’s most famous work was called Principia…”, the next word must be “Mathematica” — the watermark has nothing to act on. Rigorous technical or factual writing therefore carries a thinner signal than creative prose.
Proofreading barely registers. If Claude lightly edits a human’s text, nearly all the words are the person’s own. The watermark lives only in Claude’s changes, and a handful of grammar fixes may not register at all.
Code is the weakest case. Code that must be exact — where a different token would break it — gets no watermark nudge at all. The mark survives mainly in arbitrary choices like comments and variable names, where either option is equally valid. For teams hoping to detect AI-written code with this scheme, the realistic coverage is partial at best.
And one more boundary: the watermark answers only “what is the likelihood this was partly written by Claude?” It cannot confirm a text is human-written, and it cannot attribute text to a different AI — even another watermarked model, which would use a different key or an entirely different method.
Files Get Cryptographic Provenance, Not Just Patterns
Text is only half the marking system. When Claude generates a supported file type — .svg, .png, .jpg — it now attaches signed provenance metadata following the C2PA open standard (the same Coalition for Content Provenance and Authenticity spec used across the industry). A signed label signals the file was processed by Claude and lets verifiers detect tampering after the fact.
The two mechanisms are complementary: the text watermark travels with copied-and-pasted words, while C2PA metadata anchors files cryptographically. Neither is bulletproof alone; together they cover the two dominant forms Claude’s output takes in real workflows.
The Brussels Effect, Ship Once
Legally, Article 50(2) only binds systems used within the EU. Nothing compelled Anthropic to mark a response generated for a user in Singapore or São Paulo. It did anyway, and the stated reason is disarmingly operational: the company says it doesn’t yet have a durable way to scope watermarking by region, so it shipped one global behavior.
The result is a textbook instance of the Brussels Effect. A transparency rule written for one market becomes the de facto global standard for Claude output, because complying once is cheaper than maintaining two code paths and building a region-aware detection system. Every Claude user, everywhere, is now subject to an EU-inspired transparency regime — whether or not their jurisdiction has any equivalent rule.
There is a countervailing consideration. Anthropic notes the watermark carries no user-identifying information and cannot be traced to a person or chat. But as detection tooling rolls out — Anthropic says it will support third-party detection as the Code requires — the practical question becomes who holds the keys, and under what process marks get checked. Transparency infrastructure is also surveillance infrastructure, if you squint.
What Happens Next
Three things to watch:
- Retrofitting older models. The law includes a transition period for models launched before August 2, and Anthropic says it is working to add marking to them. Watch whether — and how quickly — Claude Opus 5, launched July 24, gets marked.
- Detection access. Anthropic promises forthcoming documentation enabling users and third parties to detect Claude’s marks. The shape of that access — public detector? API? vetted partners? — will determine whether this is real transparency or compliance theater.
- The other 189 signatories. Anthropic is early, not alone. As other Code of Practice signatories ship their own watermarking (Google’s SynthID already covers text, photos, audio, and video), interoperability becomes the next battle: a world of incompatible, keyed watermarks is better than nothing, but far from the seamless provenance layer the AI Act’s drafters imagined.
For now, the quiet landmark stands: one of the world’s most widely used writing models now provably — if statistically — signs its work. The words are unchanged. The dice are rigged.
Sources
- [1] https://www.anthropic.com/news/claude-text-watermark
- [2] https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content
- [3] https://www.euronews.com/next/2026/08/11/eu-compliance-delivered-globally-anthropic-to-watermark-claudes-output-worldwide
- [4] https://c3.unu.edu/blog/claude-ai-watermark-eu-ai-act-coverage
- [5] https://techcrunch.com/2026/08/11/anthropic-says-it-will-watermark-text-generated-by-its-ai-models/