False positives · non-native English writing flagged as AI

Why non-native English writing gets flagged as AI

If you are a non-native English speaker and your own writing has been flagged, there is a real, published reason detectors do that, and it has nothing to do with whether you wrote it.

Liang et al. (2023), published in Patterns, found AI-writing detectors misclassified 61.3% of TOEFL essays by non-native English writers as machine-generated. The pattern traces to a measurable property of the writing itself: a narrower, more predictable set of words and sentence patterns, the same property a classifier is built to catch in generated text. If your writing was flagged and English is not your first language, this is a documented pattern, not a personal failing, and it gives you something specific to hand whoever raised the concern.

Why this happens

Classifiers built to catch machine text mainly score two properties: how predictable each word is given the ones before it, a property called perplexity, and how much sentence-to-sentence variation the writing carries, called burstiness. Formal English instruction for standardized exams teaches a comparatively narrow set of safe constructions: fixed connector phrases, a smaller working vocabulary chosen for accuracy over range, and sentences often planned in advance to avoid grammar mistakes. That is sound teaching, and it produces writing with lower perplexity and less sentence-to-sentence variation than a native speaker writing off the cuff, which lands close to the statistical signature a detector is trained to flag. The writer did nothing wrong. The tool is measuring a side effect of learning a language carefully, not the presence of a machine.

Where honest writing and machine writing genuinely overlap

These are real patterns, not coincidences. Each one is something a careful human writer in this situation has good reason to do, and something a model does for entirely different reasons.

The not-only-but-also connector
This linking pattern, along with 'on the other hand' and 'in conclusion,' is explicitly taught as a standard way to open a contrast in most ESL and exam-prep curricula. It is also one of the most common connective patterns in generated text, for the unrelated reason that both are statistically safe choices.
The narrower safe vocabulary
A second-language writer often chooses a word they know is correct over one that is more precise but riskier. That narrows the range of word choices in a way that lowers the text's unpredictability, the same property a classifier reads as machine-generated.
The planned-in-advance sentence
Composing full sentences mentally before typing them, to avoid a grammar mistake, reduces the small in-the-moment variation a native speaker's writing usually carries. Less variation reads as more uniform, and uniform reads as machine-made to these tools.
The formal transition stack
Furthermore, moreover, and additionally are taught explicitly as ways to open a new paragraph in exam-prep materials. Used at the rate a curriculum teaches them, they appear far more often than in ordinary native writing, which is itself an unusual pattern to a classifier.

What actually helps

Keep the same evidence anyone would: earlier drafts, an outline, notes in your first language if that is how you plan your writing. If your institution raises a flag, you can point to Liang et al. (2023) directly: detectors misclassified 61.3% of TOEFL essays by non-native writers in that study. It gives a reviewer a documented reason to treat the flag with more caution rather than less. Ask whether the institution has a policy for non-native speakers specifically, since some do and most do not yet, and ask what evidence besides the score they are weighing. None of this changes what a tool reported. It gives the person making the decision a reason to look past the score rather than at it alone.

What the flattening looks like

On the left, honest writing with the specifics drained out of it, which is the shape that reads as machine-written. On the right, the same content with the writer's own detail restored. The voice is The Steady Hand (RCOAN): plain words about a specific fact, in place of the safer textbook connector.

Before · flattened

Furthermore, environmental protection is an important issue that affects everyone in society. Not only does pollution harm human health, but it also damages the ecosystem in serious ways. On the other hand, some people argue that economic growth should come first. Countries that invest early in renewable energy tend to reduce their fuel imports within a decade. In conclusion, governments should balance development with environmental responsibility for future generations.

After · specifics restored

Environmental protection matters to everyone. Pollution harms human health and damages the ecosystem. Some people argue that economic growth should come first. Countries that invest early in renewable energy tend to reduce their fuel imports within a decade. Governments should balance development with environmental responsibility for future generations.

Steps worth taking

  1. Keep your drafting evidence. Outlines, earlier versions, and any notes in your first language are real evidence of process, whatever language they are written in.
  2. Cite the published research if it comes up. Liang et al. (2023) in Patterns, on detector bias against non-native writers, found detectors misclassified 61.3% of TOEFL essays. Naming it gives a reviewer something concrete to weigh.
  3. Ask what evidence the reviewer is using beyond the score. A single number from one tool is a weak basis for a decision on its own. Ask what else is being considered before you argue with the score itself.
  4. Ask about institutional policy. Some schools have written guidance for handling detector flags on writing from non-native speakers. If yours does not, asking is reasonable and puts the question on record.

Questions

Do AI detectors have a bias against non-native English speakers?

Yes, and it is documented. Liang et al. (2023) found AI-writing detectors misclassified 61.3% of TOEFL essays by non-native writers.

Why is my non-AI writing being detected as AI?

Because detectors score predictability and sentence variation, not authorship. Writing shaped by careful, formal language instruction can score close to generated text even though a person wrote every word.

What do I do if my own writing is being flagged as AI?

Keep your drafts and notes, name the published research on detector bias if it is relevant to your situation, and ask the reviewer what evidence besides the score they are using.

How often do students get flagged for AI?

There is no single reliable figure, and false-positive rates vary a lot by tool and by the kind of writing being scored. Liang et al. (2023) put the misclassification rate for non-native TOEFL essays at 61.3%.

underslop edits drafts for voice. It does not check, score, or contest anything, and nothing on this page will change a result you have already been given. What it can do is show you which patterns in a piece of writing read as machine-written, using the same open lint the rest of the site runs on.

See which patterns your writing actually carries