Why the em dash became the AI writing tell
By Claude Watermark Research · Updated
Large language models emit em dashes far more often than most human writers because their training data over-represents edited, professional prose where the em dash is common, and because the token is a low-risk way to join clauses. It is the single most-cited tell in accusations of AI writing — but it is weak evidence on its own, since plenty of people have always written this way.
Why models reach for it
An em dash lets a model extend a sentence without committing to the grammar a colon, semicolon or full stop would require. When the next-token distribution is uncertain, a dash keeps more continuations available. That makes it structurally attractive at exactly the moments a model is least sure.
Training data amplifies this. Published and edited prose — journalism, essays, marketing copy — uses the em dash more heavily than casual writing, and that is disproportionately what these models learned from.
Why it is weak evidence
The em dash is a legitimate, long-standing punctuation mark. Writers who came up through editing, or who read a lot of long-form journalism, use it constantly and always have. Accusations resting on dash frequency alone routinely land on people who simply write like that.
It is also trivially removable, which cuts against it as a detector: anyone actually trying to hide AI involvement would strip it in seconds. What it really flags is unedited output, not AI output.
Replacing them without flattening the prose
A blind find-and-replace to commas produces run-on sentences that read worse than the dashes did. The fix depends on the job the dash was doing: an aside becomes a pair of commas or brackets, an explanation becomes a colon, an abrupt turn becomes a new sentence.
Our cleaner replaces em dashes by default and picks the substitution from surrounding structure rather than swapping every one for the same character. Frequency is worth checking in your own writing too — if you use several per paragraph, that is a style note independent of any AI question.
Questions
- Do em dashes prove text was written by AI?
- No. They correlate with unedited model output, but many human writers use them heavily. On their own they are not evidence of anything.
- Should I remove every em dash?
- Only if the surrounding sentence still reads well. Removing them mechanically usually produces worse prose than leaving them in.
- Do detectors look for em dashes?
- Statistical detectors look at overall token distributions, of which punctuation is one small part. Human accusers are the ones fixating on the dash specifically.
Check your own text. Free, unlimited, no account, and it runs in your browser so nothing is uploaded.
Open the checker






