Claude Now Writes With an Invisible Watermark: Inside Anthropic's SynthID Implementation
Anthropic's new Claude models embed an undetectable SynthID-based watermark into every generated word — here is exactly how the key-driven scheme works, what it can and cannot prove, and why it is rolling out globally.
Every word that new Claude models write now carries a signature no reader will ever see. On August 14, 2026, Anthropic published a detailed technical explainer confirming that its latest models weave an invisible, machine-readable watermark directly into generated text — a change that took effect as the EU AI Act’s transparency obligations became binding on August 2. The move makes Anthropic one of the first major providers to ship text watermarking at scale, and it lands in the middle of a fierce global debate over how to police AI-generated content without breaking the open internet.
What changed, exactly
As of the August 2 compliance deadline, AI providers serving the European market must mark machine-generated content under Article 50 of the EU AI Act. Anthropic, along with roughly 190 total signatories including several other major model developers, signed the EU Code of Practice on Transparency of AI-Generated Content in July 2026. The code requires providers to use “marking” methods that make synthetic text identifiable after the fact.
Anthropic’s answer is a cryptographic watermark embedded in the text itself. The company confirmed six properties that define the design:
- Zero visible impact. Watermarked and unwatermarked text are indistinguishable to readers. Nothing is added, and there are no hidden characters or zero-width tricks.
- No quality cost. Internal testing found no measurable impact on content, creativity, or readability.
- No price or latency cost. The scheme produces no extra tokens, so models cost the same to serve and run at essentially the same speed.
- No identity leakage. The watermark carries no information traceable to a person, organization, or specific chat.
- Not Claude-specific by nature. The method is a shared industry approach; other providers signing the code will deploy their own keys and variants.
- Global at launch. Anthropic is applying the watermark worldwide because it does not yet have a durable way to scope marking by region — a detail that has already drawn attention outside Europe.
How the watermark actually works
The technique is a version of SynthID-Text, the scheme Google DeepMind introduced in a Nature paper in October 2024. It belongs to a family of approaches dating back to a 2022 proposal by cryptographer Scott Aaronson, and all of them share one elegant design principle: the watermark changes only the source of randomness, not the words the model considers.
Large language models generate one word at a time. At each step, the model ranks candidate next-words. Sometimes several candidates are effectively interchangeable — after “The weather today was cold and…”, the next word could plausibly be “overcast” or “grey” without changing the meaning. Normally, a random number generator settles which synonym gets picked. Watermarking swaps that arbitrary RNG for a keyed pseudorandom source: the key plus a few preceding words decide the choice.
The result reads exactly like normal text — sometimes “overcast”, sometimes “grey”, exactly as before. But anyone holding the key can scan a sequence of choices and compute the statistical likelihood that the text is consistent with Claude’s keyed process. Anthropic’s own analogy is a Monopoly game where die rolls are replaced by consecutive digits of pi drawn from a random offset: the moves look random to the players, yet the full move history betrays its source to anyone who knows pi.
Crucially, the watermark never pushes Claude toward words it wouldn’t already consider. It would not make the model choose “nubilous” — an obscure synonym for overcast — because the nudge only applies between genuinely interchangeable options.
Where the watermark lives — and where it fades
The scheme’s honesty about its own limitations is one of the more refreshing parts of the announcement:
- Factual text is sparsely marked. After “Isaac Newton’s most famous work was called Principia…”, only one next word is correct. Where there is no real choice, the watermark has nothing to act on.
- Code is lightly marked. Code that must be exact carries almost no watermark; arbitrary choices within comments can be marked, with negligible effect on the code itself.
- Light edits escape detection. If Claude merely proofreads human text, the few corrected words may not provide enough signal to register.
- Short samples are weak. Detection confidence grows with length, because longer passages contain more keyed choices.
- Rewriting kills it. Light editing probably won’t fully strip the watermark; replacing every word will. At that point, Anthropic argues, the text is arguably no longer AI-generated anyway.
- Proof of involvement, not authorship. A positive detection cannot distinguish “Claude wrote this” from “Claude heavily edited this,” and it says nothing about ownership or legal responsibility. It also cannot identify text from other providers — each uses its own key and possibly an entirely different scheme.
For files, Claude takes a different route: supported formats such as .png, .jpg, and .svg now attach C2PA content credentials — cryptographically signed metadata stating the file was made or processed with Claude. Unlike the watermark, this is not embedded in pixels; it is an open industry standard readable by any C2PA-aware tool.
A watermark detection API is in the works, Anthropic says, with implementation details still being finalized. Older Claude models launched before August 2 are covered by the law’s transition period and will gain watermarking in a rollout “over the coming months.”
Why this matters
The rollout is the first real-world, at-scale test of Aaronson-style cryptographic watermarking on a frontier consumer product, and its stakes cut three ways.
For regulators, it is proof-of-life for Article 50: the AI Act’s transparency regime can technically work without visible labels on every paragraph. For the industry, it sets a coordination precedent — 190 signatories committing to interoperable-in-spirit marking, with Google’s SynthID lineage already embedded in Gemini. And for users, it draws the first crisp line between provenance (who or what made this) and detection (guessing from stylistic tells). Anthropic explicitly contrasted its keyed watermark with statistical detectors like Pangram, which hunt for phrasing habits — the AI fondness for “this isn’t X, it’s Y” constructions and conspicuous use of “quietly.”
The open questions are just as significant. Security researchers note that keyed watermarks survive copy-paste but fall to determined paraphrasing, and Google’s own SynthID has been broken by developers before. The global rollout means non-EU users are getting watermarked output under a European law they had no say in — a preview of the Brussels-effect friction to come. And the detection API’s access model matters enormously: whether it lands open to platforms, restricted to vetted parties, or gated by courts’ demands will shape whether watermarking becomes genuine transparency infrastructure or a surveillance tool by another name.
One thing is certain: the era of unverifiable AI text is ending, one statistically loaded synonym at a time.
Sources
- [1] https://www.anthropic.com/news/claude-text-watermark
- [2] https://www.nature.com/articles/s41586-024-08025-4
- [3] https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content
- [4] https://www.cooley.com/news/insight/2026/2026-08-03-eu-ai-act-transparency-obligations-take-effect-2-august-2026
- [5] https://techcrunch.com/2026/08/11/anthropic-says-it-will-watermark-text-generated-by-its-ai-models/