anthropic.com web signal

Anthropic Details Claude Watermark's Gaps on Code and Rewrites

Anthropic AI Detection ai-business

TL;DR

  • Claude will ship an invisible watermark based on Google DeepMind's SynthID-Text, added to help Anthropic comply with the EU AI Act.
  • The mark barely covers code and other outputs that must be exact, because there are no interchangeable synonyms for the technique to swap.
  • Proofreading a human draft or fully rewriting Claude's output removes the signal, and detection also weakens on short samples.

Anthropic has published a technical explainer of the invisible text watermark it is rolling into future Claude models, and the interesting part is how much of Claude's actual output it does not cover. The technique is a version of Google DeepMind's SynthID-Text, described in a 2024 Nature paper, and it works by nudging Claude toward one of several equally valid word choices at low-stakes moments, using a key plus the preceding words to decide which synonym to pick. The company says the change is going in to help comply with the EU AI Act, which our tracker followed when Anthropic signed the bloc's code of practice.

The pitch is that quality does not move. Anthropic says the difference between watermarked and unwatermarked text will not be distinguishable to readers, that watermarking does not require extra tokens and will not be more expensive, and that speed impact is negligible. It cites Google DeepMind's own live test on a portion of Gemini traffic, which reportedly found no statistically significant differences from the unwatermarked model.

The caveats are the point. Because the mark lives in choices between interchangeable words, it barely shows up when the output has to be exact. Code, Anthropic writes, 'has generally less watermarking than some other forms of text.' Proofreading is another blind spot: when Claude only lightly edits a human draft, 'nearly all the words are the person's, there's very little (if anything) for the watermark to attach to.' A complete rewrite where every word is replaced strips the signal entirely, and detection also weakens on short samples.

Anthropic is unusually blunt about what the tool can and cannot prove: 'A watermark can only determine that Claude was likely involved with the content at some point. It cannot distinguish "Claude wrote this" from "Claude heavily edited this."' The post does not publish a false-positive rate, a minimum reliable sample length, or say who outside Anthropic gets the detector key, so anyone building policy on top of this signal is working with an incomplete spec.

For educators, publishers and platforms trying to label AI-assisted writing, this offers a real but partial provenance signal, landing the same week Google let Gemini and Flow users toggle off visible watermarks and adding to our 29 AI detection stories in the last 90 days. Anyone motivated to hide model involvement already has an obvious playbook: ask Claude to proofread, or rewrite the output somewhere else. The compliance story here is stronger than the enforcement one.

Shared on Bluesky by 2 AI experts