What a Claude text watermark detection proves
A watermark result reads like a verdict and is not one. Both directions of the inference are weaker than they look.
It shows involvement, not authorship
The strongest claim a detection supports is that Claude was likely involved with the text at some point. It cannot separate writing from heavy editing.
That distinction decides how the result can fairly be used. Someone who wrote a piece themselves and asked Claude to tighten the prose can produce a marked document, and the mark looks the same as one from a prompt that produced the whole thing. Any policy that treats a hit as proof of AI authorship is reading a claim the evidence does not make.
"A watermark can only determine that Claude was likely involved with the content at some point. It cannot distinguish \"Claude wrote this\" from \"Claude heavily edited this.\""
A clean result is not evidence of a human author
The absence of a mark carries almost no information. The key answers one question — how likely it is that Claude was partly involved — and that question has no bearing on whether a person or a different model produced the text.
Two ordinary situations produce unmarked AI text: output from a model other than Claude, and text that has been rewritten from scratch. A clean result is consistent with both of them, and with genuine human writing.
"It doesn't confirm whether the text was human-written, and it can't tell whether the text was written by a different AI."
Where the mark thins out
The watermark rides on choices between near-equivalent words. Where the writing has few such choices, there is little left to carry it — which means the signal is weakest exactly where text is most consequential.
Short passages carry too little to read
Detection needs volume. A short sample offers fewer word choices, so there is less information to work from and the result is correspondingly unreliable.
This is a property of the method rather than a bug to be fixed later. A paragraph in an email, a headline, a single answer in a form: these are the units most people would want checked, and they are the units the method handles worst.
"Detecting a watermark also doesn't work well on small samples, where there are fewer word choices and thus less information to go on."
Facts and code leave the least room
Accuracy and watermarking pull in opposite directions. Where only one wording is correct, there is no free choice to encode anything into, so factual passages carry a sparser mark than discursive prose.
Code is the extreme case: it usually has to be exact, so it carries generally less watermarking than other kinds of text. The practical consequence is uncomfortable. Reference material, technical documentation and source code — the categories where provenance matters most — are the categories where the mark is thinnest.
"code—which in very many cases has to be exact—has generally less watermarking than some other forms of text."
Editing weakens it in proportion to how much you change
Light editing probably will not remove the mark; replacing every word will. Everything else sits on the line between those two ends, in proportion to how much of the original wording survives.
There is no threshold published, and there could not usefully be one, because the mark is statistical rather than a stamp in a fixed location. Rewriting a marked draft in your own words does not partially erase a tag — it removes the word choices the signal was made of.
"To some extent, yes. Light editing probably won't remove the watermark completely; a complete rewrite where every word is replaced will."
What it carries, and what you can do with it today
Two questions decide whether this matters in practice: what the mark reveals about you, and whether anyone outside Anthropic can read it.
No identity is encoded in the mark
The watermark carries no identifying information and cannot be traced to a specific person, organization or chat. It marks the output, not the account that produced it.
This is worth stating plainly because the intuition runs the other way — a hidden mark sounds like a tracking tag. The key detects a statistical pattern in word choice. There is no field in it for a user, and nothing in the key that would let anyone recover one.
"Watermarking carries no identifying information and can't be traced to a specific person, organization, or chat."
Images get content credentials instead
Files take a different route entirely. When Claude produces a supported file type, it attaches a content credential — a small, cryptographically signed note in the file's metadata.
The trade-off is the mirror image of the text case. A signature in metadata is precise and verifiable, and it survives nothing: a screenshot, a re-export or a format conversion drops it. Text watermarking is vague but travels with the words; content credentials are exact but travel only with the file.
"When Claude produces a file of a supported type (such as a .png, .jpg, or .svg), it will attach a content credential in the form of a small, cryptographically signed note in the file's metadata."
Nobody outside Anthropic can check a document yet
Reading the mark requires the key, and the key is Anthropic's. A detection API is promised, with implementation details still being worked out.
Until it ships, the practical position for a school, publisher or employer is that this changes nothing about what they can verify themselves. Third-party AI detectors are unaffected by any of this: they do not have the key and are inferring from style, which is a different method with its own error rate.
"We will soon be offering a watermark detection API. We're in the process of working out the details of its implementation."
Anthropic's announcement is a page of tables and nested FAQ entries, and the limits are stated in the answers rather than the headlines. Converting the page to markdown keeps those answers attached to their questions, which is the difference between reading the policy and reading the summary of it.
The honest summary is narrow. A detected watermark means Claude probably touched the text; an undetected one means nothing much at all; and the signal is weakest on short, factual and technical writing. That is a useful signal for measuring how much AI-written text is circulating at scale. It is a poor basis for a decision about one document and one person — and until the detection API ships, it is not a decision anyone outside Anthropic can make anyway.



