Claude · The skeptic's guide

Claude and AI detectors: what they actually flag

A detector's verdict on Claude text says nothing about Claude. It's a statistical guess about how predictable the text is.

AI detectors cannot tell you that text came from Claude. They are statistical classifiers. They score how predictable each word is, given the words before it. They do not read watermarks, and they do not read C2PA metadata. So Claude text flags for style reasons, not because a mark was found. A flag is a probability, not a finding. OpenAI's own classifier was withdrawn after a 26% true-positive rate and a 9% false-positive rate. Studies found detectors flag more than half of non-native English essays as AI. University libraries call detectors 'problematic and not recommended as a sole indicator' of misconduct. The vendors themselves advise against acting on a score alone. Anthropic's detection tooling for the new marks is still forthcoming.

Why this page exists

The results mix detector tools with dated claims like 'Claude doesn't show up on detectors' from 2023. Searchers want to know whether their Claude-assisted work will be flagged. The honest answer: what the tools can and cannot measure, plus the real false-positive stakes.

The tells that show up here

Not every AI tell appears everywhere. These are the ones that cluster on this surface, and what it is about the format that produces them.

vague-attribution
'Some detectors say' is the habit to replace with named, dated findings.
not-x-its-y
The page's core distinction (flag vs verdict) must avoid the parallel tic.
em-dash-chain
Detector reviews are dense with dash chains. Keep this page's copy clean.

Before and after

A generic AI draft on the left. On the right, the real output of running that draft through underslop, checked so it adds no name or number the draft did not already carry. The voice is The Steady Hand.

Before · generic AI draft

As an AI language model, I'm often asked: can detectors catch Claude's writing style? It's not just a question — it's a testament to how far AI detection has come! Research consistently shows that detectors can identify Claude's distinctive voice with near-perfect accuracy. Here's the thing, though: results vary, so check with multiple tools to be sure!

After · edited

Detectors can often catch Claude's writing style. Their scores flag the patterns, not the author. Results vary between tools. A single score is a signal, not a verdict. Check the tells yourself before you act on it.

How to do it by hand

  1. Treat the score as a number, not a verdict A flag is a probability estimate, not a finding of fact. OpenAI's own classifier mislabeled 9% of human text. Nothing is decided until a person acts on the score.
  2. Know the tool's limits Detectors score how predictable text is. They don't read watermarks or C2PA metadata. Universities have turned off detector features, and libraries call the tools problematic as a sole indicator.
  3. Check the actual tells in the draft Look for the structural patterns that push a predictability score higher: validation openers, dense dashes, parallel constructions. If they show up, the flag is about style, not about a mark.
  4. Keep evidence and talk to the instructor Save drafts and version history before you write. Ask what was flagged and by which tool. The score isn't proof. Cases have turned on due process, not the number.

Questions

Can an AI detector tell if text came from Claude?

No. Detectors are statistical classifiers. They don't read watermarks or provenance metadata, and they can't attribute text to a specific model.

Does a detector flag Claude text?

Sometimes it's based on statistical patterns, not marks. The rate changes with the tool and the text. No detector reads the Claude watermark, C6 or C50.

Is Claude easier or harder to detect than other models?

Detectors cannot identify which model wrote a text, so there is no verified model-level comparison. The 2023-era claim that Claude "doesn't show up" on detectors no longer holds. Detection tooling for the new marks is still in the works.

Can a detector flag human writing?

Yes, at measurable rates. OpenAI's classifier mislabeled 9% of human text, and studies found detectors flag over half of non-native English essays.

Are detector scores proof of anything?

No. University libraries call detectors "problematic," and vendors say a score should not be the only reason for action. They advise using it with other checks, not alone.

Claude marks are real and the coverage is still settling. underslop is the edit, not a remover: it edits an AI draft so it reads like you wrote it, and it makes no promise about what any detector will report. Your text is held for 0 seconds.

Why Claude text gets flagged as AI: not always a watermark