AI Guides › Playbooks
By Nigel Guy · 8 min read
Most people who hear that Claude now watermarks its text react in one of two ways. Some panic that every client email is "tagged". Others shrug it off because they can't see anything. Both lead to the same mistake: you keep making claims about your work, either "this is all mine" or "nobody can tell", that you never checked against what the mark does. The mark is real, but it is narrower than the panic suggests and harder to dodge than the shrug assumes.
The rule: decide what you will claim about a piece of work before you send it, and base that claim on how the work was made, never on whether you think a detector would catch it.
Anthropic's support centre says that supported Claude models now add an imperceptible watermark to the text they generate. It is not hidden characters, it does not use extra tokens, and Anthropic says it carries no identifying information. It cannot be traced back to you, your organisation or a particular chat.
The mechanism, as Anthropic's technical post describes it, works at the point where Claude chooses between words that mean roughly the same thing. That choice normally comes from a random number. With watermarking, it comes from a secret key combined with the few words just before. Someone holding the key can test a passage for the pattern. Anthropic says the method builds on SynthID-Text, which Google DeepMind published in Nature in 2024.
Files are handled separately. When Claude produces a supported file, such as a PNG or JPEG, it attaches signed Content Credentials metadata using the C2PA open standard.
Article 50 of the EU AI Act applies from 2 August 2026. Paragraph 2 requires providers of systems that generate synthetic audio, images, video or text to mark the output in a machine-readable form that can be detected as AI-generated. Anthropic has signed the Code of Practice on transparency of AI-generated content that supports that article, and says models launched on or after 2 August support marking from day one.
Three details are worth knowing:
| Output | Marked? | Strength of the mark |
|---|---|---|
| Text Claude writes from scratch (Claude apps, Claude Code, Cowork, Claude Tag, the API) | Yes | Strongest on long, discursive prose |
| Translations by Claude | Yes | Anthropic notes that Claude chooses every word |
| Claude proofreading or lightly editing your text | Barely | Almost all the words are still yours |
| Factual passages, names, figures | Sparse | Few safe word choices to work with |
| Code | Negligible | Syntax leaves no room; comments may carry it |
| Short snippets (a subject line, a tagline) | Hard to detect | Too few choices to test |
| PNG or JPEG files Claude generates | C2PA metadata | Removed if metadata is stripped |
| Claude via AWS, Google Cloud or Microsoft Foundry | Text, yes | C2PA only where the platform supports file generation |
Who can check? For files, anyone can use Anthropic's free Claude Content Checker. It reads the credentials on your device and does not upload the file. For text there is a detection API, which at time of writing is in private preview. It is open to regulators, law enforcement, media, fact-checkers, independent researchers, educational organisations, EU civil society groups and enterprises with compliance obligations. You probably can't test your own text yet; a newsroom or regulator might.
The trap is treating the watermark as the thing that decides what you can claim. Anthropic is explicit on both sides:
So "it won't be detected" is not evidence that a claim is true. "I only used it for a light edit" is not protection either. If you pasted in a full Claude draft and changed a few words, the mark probably survived, and Anthropic says light editing probably won't remove it completely.
Fill this in for anything that goes out under your name: a report, a client deliverable, a published article, coursework, a submission.
A freelance consultant in Leeds writes a 12-page market summary for a client who sells into Germany. Claude drafted sections 2 and 3. She rewrote the conclusions and checked every figure against the original sources. That is L2. Her claim is "prepared by me, with AI drafting assistance on the background sections". She keeps the chat export alongside her final draft. She does not tell herself that the rewrite "probably killed the watermark". Whether it did or not changes nothing about what is true.
You are helping me write an honest, plain-English note about how AI was used on one piece of work.
Context:
- The piece: [WHAT_IT_IS_AND_LENGTH]
- Who will read or rely on it: [AUDIENCE]
- Any rules the audience has set (contract clause, publisher policy, academic regulation): [RULES_OR_NONE]
- Exactly what the AI did, section by section: [WHAT_AI_DID]
- Exactly what I did myself, including fact-checking: [WHAT_I_DID]
Steps:
1. Put the work at one level: L0 (no AI), L1 (AI proofread only), L2 (AI drafted parts, I substantially rewrote), L3 (AI drafted most, I edited), L4 (AI output, lightly checked). Give one sentence explaining your choice.
2. Draft a disclosure line of no more than 30 words that matches that level and meets any rules I have listed.
3. List anything in my description that would make the disclosure misleading.
Constraints: do not suggest wording intended to hide AI involvement or to avoid detection. Do not tell me whether a watermark would be detected. If any bracketed input is missing or vague, ask me for it before you draft.
Before answering, check that the level you chose matches what I said the AI did, not what I would like it to be.
Fill in the five bracketed fields with what actually happened, not what sounds best.