ESL Humanizer

Why AI detectors flag human writing — and why non-native writers get hit hardest

By ESL Humanizer editorial teamUpdated 8 min read

Short answer

Most AI detectors score how predictable your wording is. Non-native writers tend to use common words and simple sentence structures, which look predictable — so their genuine work is flagged far more often. In a Stanford study, seven detectors labeled 61% of human-written TOEFL essays as AI-generated on average, while classifying essays by US eighth-graders almost perfectly.

If you wrote something yourself and an AI detector still said it was machine-written, you are not imagining a problem — and you are not alone. The way most detectors work means that careful, plain English is exactly the kind of writing they are most likely to get wrong. That hits non-native speakers hardest.

This guide explains, without jargon, how detectors reach their verdicts, what peer-reviewed research says about their accuracy for ESL writers, and what you can reasonably do about it.

How AI detectors decide whether text is “AI”

AI detectors do not know who wrote a text. They cannot see your screen, your drafts or your browser history. They only see the words, and they estimate how likely those words are to have come from a language model.

Most detectors lean heavily on a measure called perplexity — roughly, how surprising each next word is. Language models like ChatGPT tend to choose likely, common words, so their output has low perplexity. Detectors therefore treat low perplexity as a sign of AI. Many also look at burstiness: human writers usually mix long and short sentences, while model output is more uniform.

The flaw is obvious once you see it: humans also write predictable text. Anyone writing carefully in a second language — choosing safe words, avoiding idioms they are unsure of, keeping sentences regular — produces exactly the pattern detectors associate with machines.

Why ESL writing gets flagged more often

Stanford researchers who studied this problem put it directly: detectors “typically score based on a metric known as ‘perplexity,’ which correlates with the sophistication of the writing” — something in which non-native speakers naturally trail native speakers (Stanford HAI).

Non-native writers tend to score lower on the features that raise perplexity:

  • Lexical richness and diversity — using a smaller set of common words rather than rare synonyms.
  • Syntactic complexity — shorter, more regular sentences built on familiar patterns.
  • Grammatical complexity — fewer embedded clauses and less variation in structure.

None of these are signs of cheating. They are signs of someone writing clearly in a language they are still mastering. But to a perplexity-based detector, they look the same as AI.

What the research found

The most cited study on this question, published in the journal Patterns in 2023, ran seven widely used AI detectors on two sets of human-written essays: 91 TOEFL essays written by non-native English speakers, and essays written by US eighth-graders (Liang et al., 2023).

FindingResult
US eighth-grade essaysClassified “near-perfectly” as human
TOEFL essays misclassified as AI (average across 7 detectors)61.22%
TOEFL essays flagged as AI by all seven detectors18 of 91 (19%)
TOEFL essays flagged by at least one detector89 of 91 (97%)

Every one of those essays was written by a human. The researchers also showed that simple prompting — asking a model to rewrite text with more sophisticated vocabulary — both reduced the bias and let AI-generated text slip past the same detectors. In other words, detectors were penalizing limited vocabulary, not AI use. The authors cautioned against using detectors in evaluative or educational settings, particularly where they may penalize non-native speakers.

How reliable are AI detectors in general?

The ESL problem sits on top of a broader accuracy problem. Two data points from the organizations closest to the technology are worth knowing:

  • OpenAI, the company behind ChatGPT, released its own AI-text classifier in January 2023. By its own evaluation it caught only 26% of AI-written text and wrongly labeled human writing as AI 9% of the time. OpenAI withdrew it on July 20, 2023, citing its “low rate of accuracy” (OpenAI).
  • Vanderbilt University disabled Turnitin's AI detector in 2023. It noted that Turnitin had claimed a 1% false positive rate, and that with 75,000 papers submitted in 2022, around 750 student papers could have been wrongly labeled (Vanderbilt).

Even a “small” false positive rate becomes a large number of real students when a tool is run on every submission — and the Stanford data suggests the rate is much higher for ESL writers than for native speakers.

What this means for you

If you write in English as a second language, three practical conclusions follow:

  1. Keep evidence of your process. Write in Google Docs or Word with version history on, save drafts and notes, and keep your sources. This is the single best protection against a false accusation. See what to do if you are accused.
  2. Don't “fix” your writing by making it sound fancier. Stuffing in rare words can make your work sound less like you and invite different questions. Natural, idiomatic English — the way a fluent writer would phrase the same idea — is the goal.
  3. Know the difference between polishing and disguising. Improving the English of your own ideas is not the same as passing off AI text as yours. Our guide on whether using a humanizer is cheating explains where the line is.

See how natural your writing reads

ESL Humanizer rewrites your own paragraph into natural English, keeps your meaning and level, and shows every change word by word.

Try it free

Quick answers

Can an AI detector prove I used AI?

No. Detectors produce a probability estimate based on word patterns. When OpenAI released its own classifier, it warned that it “should not be used as a primary decision-making tool.”

Why did my essay get flagged when I wrote every word?

Most likely because it uses common vocabulary and regular sentence structures — the same low-perplexity pattern detectors associate with AI. That is typical of careful second-language writing.

Will writing longer texts help?

Longer texts give detectors more to work with and OpenAI noted its classifier was especially unreliable on texts under 1,000 characters. But length does not fix the underlying bias against simpler English.

Should I run my work through an AI detector before submitting?

You can, but different detectors disagree with each other, so a “pass” on one tool guarantees nothing on another. Keeping evidence of your writing process is more useful.

Sources

  1. Liang, Yuksekgonul, Mao, Wu & Zou — Patterns (Cell Press) (2023). GPT detectors are biased against non-native English writers
  2. Stanford Institute for Human-Centered AI (HAI) (2023). AI-Detectors Biased Against Non-Native English Writers
  3. OpenAI (2023). New AI classifier for indicating AI-written text
  4. Vanderbilt University (2023). Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector