Claude · The false-positive explainer
Why Claude text gets flagged as AI: not always a watermark
The flag is real. Usually it is not the watermark that gives it away. Here is how the mechanism works.
Most flags on Claude text are statistical judgments about style, not watermark detections. Detectors score how predictable each word is. The more a text follows the expected shape, the higher the score. Claude drafts are predictable in four big ways. An opener announces and validates the question. Hedges stack on hedges. Lists scaffold the point with headers and bullets. Em dashes chain clauses into one line. Each of those is a style reflex, not a mark. Anthropic's watermark may not even be in the text. A detected mark is not conclusive, and the lack of a mark does not prove the text was human.
Why this page exists
Students and professionals whose flagged work was flagged assume the watermark is the cause. It isn't. The statistical mechanism sits apart from the watermark. That separation is the honest answer, and it links to the existing false-positive family.
The tells that show up here
Not every AI tell appears everywhere. These are the ones that cluster on this surface, and what it is about the format that produces them.
- chatbot-artifact
- Validation openers the model adds. Readers and detectors both key on them.
- em-dash-chain
- The formatting density that makes machine text look predictable.
- not-only-but
- Parallel structure that inflates the predictability score.
- throat-clearing
- 'Let's dive in' pivots that cost the draft its voice.
Before and after
A generic AI draft on the left. On the right, the real output of running that draft through underslop, checked so it adds no name or number the draft did not already carry. The voice is The Steady Hand.
Great question — and it's one students are asking more every day! Let's dive in. It's not just the watermark, but the way the text itself is assembled: the opener that validates before it answers, the hedges piled on hedges — 'seems to suggest', 'may indicate', 'it is worth noting' — the bold-header lists that scaffold every section, and the em dashes that chain what should be separate sentences into one unbroken line — a testament to how far detection has come in today's ever-evolving landscape. Research consistently shows that the more predictable the pattern, the higher the score.
Most flags come from style, not from a watermark. Detectors score how predictable the text is, and predictable text reads as machine-made even when a person wrote it. Four patterns do most of the work. Openers announce the answer before giving it. Hedges pile up three deep. Lists scaffold every paragraph. Em dashes pack two sentences into one line. Cut those first, then keep the facts. A watermark may not even be in the text. A detected mark is not proof, and the absence of one proves nothing either way.
How to do it by hand
- Cut the opener The sentence that announces or validates before answering adds nothing.
- Say it once Hedges stack because each one softens the last. Keep the single qualifier you actually mean, like "probably" or "I think," and cut the rest.
- Break the scaffold A run of bullets and same-length sentences is exactly what detectors score. Keep a list only where it earns its place. Vary the sentence length everywhere else.
- Keep the plain word Dashes and parallel pivots are formatting, not meaning. Put the clauses back into separate sentences and state the point in the plain word. The draft keeps its facts and loses its shape.
Questions
Why does Claude text look AI-generated?
AI detectors catch a pattern, not a stamp. They measure how predictable the writing is. Claude tends to be predictable in the same ways. The patterns are a stated opener, stacked hedges, list scaffolding, and em dashes.
What makes text sound like Claude?
The tells of machine text are familiar. An opener validates before it answers. Parallel structure is performed rather than earned. Clusters of dashes appear. Bullet lists scaffold the point. Picking these out is statistical, not a watermark match.
Is it a watermark that flags my text?
Almost never directly. Detectors score predictability, not watermarks. Anthropic hasn't published its detection tooling yet, and a positive read isn't conclusive anyway.
Can human writing be flagged?
Yes, at measurable rates. OpenAI's own classifier mislabeled 9% of human text. Studies found detectors flagged over half of non-native English essays as AI.
Is a detector flag a verdict?
No. A flag is a score. University libraries call detectors problematic as a sole indicator. Universities have disabled the features. Academic cases have turned on due process, not on the score.
Claude marks are real and the coverage is still settling. underslop is the edit, not a remover: it edits an AI draft so it reads like you wrote it, and it makes no promise about what any detector will report. Your text is held for 0 seconds.