AI watermarks vs AI detectors: not the same thing
By Claude Watermark Research · Updated
A watermark is added by the model that generated the text and is verified with a secret key, so it is precise but only works on marked models. A detector like Turnitin or GPTZero is a statistical classifier applied after the fact by a third party, with no key and no ground truth — which is why detectors produce false accusations and watermarks, in principle, do not.
Why detectors misfire
Detectors infer authorship from surface statistics such as perplexity and burstiness. Non-native English writers, technical writing, and heavily edited prose all read as low-variance, so they get flagged. This is the mechanism behind the steady stream of wrongly-accused students.
No detector can prove authorship. A score is a probability estimate presented with unearned confidence.
Why watermarks are different
A watermark is a signal the generator deliberately inserted, checked against a key. When present it is strong evidence of processing by that model. But absence proves nothing, short passages carry no reliable signal, and any rewriting degrades it.
Anthropic itself says a detected mark is “not fully conclusive” — because text a human wrote and Claude merely proofread carries the same mark as text Claude authored outright.
Check your own text. Free, unlimited, no account, and it runs in your browser so nothing is uploaded.
Open the checker