NewAnthropic is watermarking Claude’s text output from 2 August 2026.See which models carry it →

AI watermark detectors: what actually exists

By Claude Watermark Research · Updated

There is no public detector for Claude's statistical text watermark. Anthropic's detection API is now in private preview for eligible organisations such as regulators, researchers, educational organisations and enterprises with compliance obligations. Public tools marketed as watermark detectors are therefore doing one of two other things: finding copy-paste artifacts such as HTML class names and invisible characters, which is real and verifiable, or running a style classifier like GPTZero, which returns a probability rather than reading Anthropic's mark. Knowing which of the three you are being shown is the entire subject.

Check and remove AI traces from my text. Checking runs in your browser, no account, nothing uploaded.

The three different things sold under one name

Watermark detection. Verifying a keyed statistical mark embedded during generation requires the vendor's key. Anthropic now offers its detector in private preview to eligible organisations, but no public checker can call it. Google publishes SynthID-Text research without a public verifier for arbitrary text. Treat any public tool claiming verified Claude watermark detection as describing access it does not have.

Artifact checking. Finding what a chat interface physically left in the text — provider HTML class names, `data-` attributes, zero-width characters. Completely reliable, because it is reading bytes rather than inferring. This is what most so-called detectors actually do.

AI-style classification. GPTZero, Pangram, Turnitin and similar. They score how predictable and uniform the writing is and return a probability. Real products with real research behind them, but they judge style, not provenance, and they are wrong in both directions.

How to tell which one you are looking at

Read what it reports. Counts and positions — *4 zero-width characters, 2 class attributes* — mean artifact checking, and the result is a fact. A percentage means classification, and the number is a guess. A confident yes or no about a vendor's watermark means the tool is describing a check nobody can currently run.

The strongest tell is whether a tool states its limits. We audited six tools in the Claude watermark category by reading their own marketing verbatim: two overclaim outright, three state their limits correctly, and the most honest documentation belongs to a free open-source repository.

What you can verify yourself, today

Artifacts, completely. If text carries `class="font-claude-response-body"` or a zero-width joiner, that is in the bytes and any checker will find it, including free ones.

We measured what this means in practice. Across roughly 50,000 words of human writing spanning 1813 to 2026 — Austen, Melville, Dickens, Joyce, an IETF specification and Wikipedia — zero deterministic artifacts appeared. Those classes really do indicate text passed through a chat interface.

The stylistic classes are the opposite, and this is where false accusations come from. Melville uses 87 em dashes in 3,400 words of Moby Dick; Austen and Bram Stoker use none at all. A signal ranging from 0 to 26 per thousand words across canonical human authors has no baseline to accuse anyone against.

One class we had to correct ourselves on: non-breaking spaces are byte-exact but appear in ordinary web copy constantly, because HTML ` ` is standard typography. We measured them on bbc.com, gov.uk and smashingmagazine.com, and fixed our own tool's wording when it implied otherwise.

What Anthropic's private-preview detector will and will not prove

Anthropic's detector is now in private preview, and its stated limits matter more than the access change. Anthropic says a watermark can only indicate that Claude was likely involved with content at some point; it does not establish authorship.

That is a provenance signal, not an authorship verdict, and the distinction is the whole thing. A detection will not establish who wrote a document, and a teacher or employer treating a positive result as proof of cheating would be reading it as something its own vendor says it is not.

It also will not work on everything. Anthropic notes the mark is weaker on short samples and on factual passages where word choice is constrained — there are simply fewer decisions available to encode it in. And it is a Claude detector: it says nothing about GPT, Gemini or Grok output.

On removal, Anthropic is equally direct: "Light editing probably won't remove the watermark completely; a complete rewrite where every word is replaced will." We take that at face value rather than dressing it up. Stripping invisible characters does nothing to a statistical mark, and no honest tool should imply otherwise.

What to do if you have been accused

Nothing detected a watermark. No such detector was available to whoever accused you, so whatever they ran was a style classifier returning a probability.

The useful response is to ask which tool produced the score and what its published false-positive rate is, and to note that the same tools flag writing from before large language models existed. Documented process, drafts and version history are worth more than any counter-score.

Questions

Is there an AI watermark detector that works?
Not publicly. Anthropic's watermark detection API is in private preview for eligible organisations, but public tools cannot call it. Tools sold to the general public as watermark detectors are therefore artifact checkers, style classifiers, or overclaiming.
Will Anthropic's detection API prove someone used Claude?
Not in the way people expect. Anthropic's own wording is that a watermark can only show Claude was likely involved at some point, and that it cannot distinguish Claude writing something from Claude heavily editing it. It is a provenance signal rather than proof of authorship, it is weaker on short or factual text, and it only covers Claude.
Can anyone detect Google's SynthID in text?
Not as a public tool for arbitrary text. SynthID-Text is published research and a reference implementation exists, but detection needs the key used at generation. The same limitation applies to Anthropic's Claude watermark.
Are AI detectors like GPTZero the same thing?
No. They classify writing style and return a probability. They cannot read any vendor's watermark and would behave identically if no watermark existed. They also produce false positives on human writing, which is the source of most wrongful accusations.
What can I check for free?
Copy-paste artifacts — HTML class names, data attributes, zero-width and exotic Unicode characters, typography. Several free tools do this correctly, including ours, and the checking runs in your browser so nothing is uploaded.

Check your own text. Free, unlimited, no account, and it runs in your browser so nothing is uploaded.

Check and remove AI traces
Product Hunt — featuredProduct Hunt — featuredNick Launches — featuredNick Launches — featuredDang.ai — featuredDang.ai — featuredFazier — featuredFazier — featuredStartup Fame — featuredStartup Fame — featuredTurbo0 — featuredTurbo0 — featuredTinyLaunch — featuredTinyLaunch — featuredToolPilot — featuredToolPilot — featuredTwelve Tools — featuredTwelve Tools — featuredProduct Hunt — featuredProduct Hunt — featuredNick Launches — featuredNick Launches — featuredDang.ai — featuredDang.ai — featuredFazier — featuredFazier — featuredStartup Fame — featuredStartup Fame — featuredTurbo0 — featuredTurbo0 — featuredTinyLaunch — featuredTinyLaunch — featuredToolPilot — featuredToolPilot — featuredTwelve Tools — featuredTwelve Tools — featured