ESL Humanizer

AI detector false positive rates: the numbers, with sources

By ESL Humanizer editorial teamUpdated 6 min read

Short answer

AI detectors are far less accurate than their marketing suggests, and least accurate on non-native English. Seven detectors misclassified 61.22% of human-written TOEFL essays as AI (Stanford, 2023). Turnitin reports a sentence-level false positive rate of about 4%. In a test of 14 detection tools, none reached 80% accuracy. OpenAI withdrew its own classifier for low accuracy.

AI-detection scores are now used in admissions, coursework and hiring, but the evidence on how often they are wrong is scattered across papers and company blog posts. This page collects the most-cited figures in one place, each linked to its original source, so you can quote them accurately — in an appeal, a conversation with your instructor or your own research.

The numbers at a glance

FindingFigureSource
Human-written TOEFL essays misclassified as AI (average of 7 detectors)61.22%Liang et al., 2023
TOEFL essays flagged by at least one of the 7 detectors89 of 91 (97%)Liang et al., 2023
Turnitin sentence-level false positive rate~4%Turnitin, 2023
Turnitin document-level false positive rate (documents scored above 20% AI)<1%Turnitin, 2023
Detection tools reaching 80% accuracy, out of 14 tested0Weber-Wulff et al., 2023
Human text mislabeled as AI by OpenAI's own classifier9%OpenAI, 2023
AI-written text OpenAI's classifier correctly identified26%OpenAI, 2023
Papers potentially wrongly flagged per year at Vanderbilt, at a 1% false positive rate~750 of 75,000Vanderbilt, 2023

Non-native writers are flagged far more often

The most important study for ESL writers is Liang, Yuksekgonul, Mao, Wu and Zou, published in Patterns in 2023. The researchers ran 91 TOEFL essays written by human test-takers through seven widely used AI detectors. On average the detectors labeled 61.22% of them as AI-generated, and 89 of the 91 essays were flagged by at least one detector. Essays written by US eighth-graders were classified almost perfectly (Liang et al., 2023).

The authors traced the gap to perplexity — how predictable each word is to a language model. Non-native writers tend to use more common vocabulary and simpler structures, which lowers perplexity and looks machine-like. GPTZero, one of the best-known detectors, describes perplexity and burstiness as indicators it uses (GPTZero).

Turnitin's own figures

When Turnitin launched its AI indicator in 2023 it promoted a document-level false positive rate below 1%. A few weeks later its chief product officer published more detail: the under-1% figure applies to documents the tool scores above 20% AI; the sentence-level false positive rate is around 4%; and 54% of false-positive sentences sit right next to genuinely AI-written ones. Scores below 20% are now shown with an asterisk because they are less reliable (Turnitin, 2023).

Vanderbilt University disabled the feature, noting that even a 1% rate could have meant about 750 of the 75,000 papers it submitted in 2022 being wrongly flagged (Vanderbilt, 2023). For more, see what a Turnitin AI score means.

Independent tests: no tool above 80%

A team of academic-integrity researchers led by Debora Weber-Wulff tested 14 detection tools, including Turnitin and GPTZero. All scored below 80% accuracy, and only five scored above 70%. The tools leaned towards classifying text as human, and their accuracy dropped further when AI text was paraphrased or machine-translated. The authors concluded the tools are neither accurate nor reliable enough to be used as proof (Weber-Wulff et al., 2023).

OpenAI's withdrawn classifier

OpenAI, the company behind ChatGPT, released its own AI-text classifier in January 2023. It correctly identified only 26% of AI-written text and mislabeled human text as AI 9% of the time. OpenAI warned it “should not be used as a primary decision-making tool” and withdrew it on July 20, 2023 because of its low accuracy (OpenAI, 2023).

What the numbers mean for you

  • A score is not proof. Every source above — including the companies that build detectors — treats the result as an estimate.
  • Your process is your best evidence. Version history, drafts and notes show how the work was written. See what to do if you're accused.
  • Quote the research accurately. Use the figures and links on this page in an appeal letter.
  • Natural English helps. Fixing translated-sounding phrasing in your own writing removes some of the patterns detectors associate with AI — without changing what you say.

Check your own paragraph

Paste something you wrote. ESL Humanizer fixes the English, keeps your meaning, and highlights every change — free for one paragraph.

Try it free

Quick answers

What is the false positive rate of AI detectors?

It depends on the tool and the writer. Turnitin reports under 1% at the document level for documents scored above 20% AI, and about 4% at the sentence level. For non-native writers it can be dramatically higher: seven detectors flagged 61.22% of human-written TOEFL essays as AI on average in a Stanford study.

Which AI detector is the most accurate?

No independent study has found a reliably accurate one. Weber-Wulff et al. (2023) tested 14 tools, including Turnitin and GPTZero, and found all scored below 80% accuracy and only five scored above 70%.

Are AI detectors biased against non-native English speakers?

Yes, according to peer-reviewed research. In Liang et al. (2023), 89 of 91 TOEFL essays (97%) were flagged by at least one of seven detectors, while essays by US eighth-graders were classified near-perfectly.

Sources

  1. Liang, Yuksekgonul, Mao, Wu & Zou — Patterns (Cell Press) (2023). GPT detectors are biased against non-native English writers
  2. Turnitin (Annie Chechitelli, Chief Product Officer) (2023). Understanding the false positive rate for sentences of our AI writing detection capability
  3. Weber-Wulff et al. — International Journal for Educational Integrity (2023). Testing of detection tools for AI-generated text
  4. OpenAI (2023). New AI classifier for indicating AI-written text
  5. Vanderbilt University (2023). Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector
  6. GPTZero (2023). What is perplexity & burstiness for AI detection?