Skip to content

How often do AI detectors flag human writing as AI?

Updated · 8 min read · by the HidenGPT team

There is no single AI detector false positive rate. Published figures run from 0.03% in one vendor's claim to 61% in an independent study, because each was measured on different writing, at different lengths, by people with different stakes.

Three numbers frame the range. Turnitin says its rate stays under 1% for documents with more than 20% AI writing detected, about one in every 100 fully human documents. A 2023 Stanford study found seven detectors wrongly flagged 61% of 91 essays by non-native English writers on average. OpenAI withdrew its own detector in July 2023 after it labeled 9% of human text as AI-written.

Published AI detector false positive figures, by source (read 10 October 2026)
SourceWho measuredWhat was testedFalse positive figureKeep in mind
Liang et al., Patterns (2023)Stanford researchers91 TOEFL essays and 88 US eighth-grade essays, 7 detectors, March 202361.22% average on the TOEFL essays; 97.8% flagged by at least one detector; about 5% on the eighth-grade essaysSmall sample; Turnitin not included; tools as they were in March 2023
Weber-Wulff et al., International Journal for Educational Integrity (2023)An independent academic team54 documents, 14 tools including Turnitin, spring 20236 of 14 tools flagged human writing; 96% accuracy on human English text, about 20% lower after machine translationNine human documents per type; tools as they were in 2023
Chambers and Kelley, AIED conference (2025)Academic researchersAbout 60,000 Reddit posts, one older open detector based on GPT-2Under 2% in each group, but significantly more for likely-autistic authorsOne detector; social posts, not essays
OpenAI AI classifier (January 2023)The vendorA challenge set of English texts9% of human text labeled AI-written, while it caught 26% of AI textWithdrawn on 20 July 2023 for low accuracy
Turnitin AI writing FAQ (updated October 2026)The vendorMore than 700,000 academic papers written before ChatGPT, before each model updateUnder 1% at document level, for documents with more than 20% AI writing; about 4% at sentence level in a June 2023 postNeeds 300 words of prose; scores from 1% to 19% are shown only as an asterisk
Turnitin second-language test (October 2023)The vendorUp to 2,000 texts per group from open learner and essay datasetsNo significant gap between second-language and native writers at 300 words or more; both slightly above 1%A bigger gap under 300 words; the vendor's own study
Copyleaks AI detector pageThe vendorNot stated on the page we read.03% false positive rate and over 99% accuracyA vendor claim
GPTZero homepageThe vendor, citing the RAID benchmarkRAID benchmark texts1% of human texts flagged; says it trains its ESL false positive rate down to 1%The vendor's summary of a third-party benchmark

Why the published rates disagree

Each number answers a different question.

Who measured it. Vendors test their own detector, usually on the kind of text their customers submit, and set the threshold before they publish. Turnitin is open about the trade: to hold false positives near 1%, it accepts missing some AI text, so a document it scores at 50% could contain as much as 65% AI writing. Independent studies test several tools at once on sets their authors built, and those sets are often small.

What was tested. Liang and colleagues used TOEFL practice essays; Turnitin says they were all under 150 words and that its own detector refuses anything under 300. Weber-Wulff and colleagues wrote 54 documents, the human ones about 10,000 characters each, and ran them through 14 tools in spring 2023. A rate measured on one kind of writing tells you little about another.

Which level. A document-level rate counts whole papers wrongly flagged. Turnitin's sentence-level figure of about 4% means a highlighted sentence has roughly a 4% chance of being human-written, and Turnitin says just over half of those sentences sit right next to real AI text.

When. Detectors keep changing. Most independent numbers describe tools as they were in 2023, and Turnitin says it replaced its multi-model setup with a single model in July 2026.

What 1% means across a real course

A low rate still lands on real people. Take a course of 300 students who each hand in four essays they wrote themselves: 1,200 human essays. At a 1% document rate, about 12 would be flagged. At the 9% OpenAI measured for its own classifier, about 108. At the 61% average from the TOEFL study, most of them. This is arithmetic to show scale, not a measurement of any class.

Two things make it worse for some writers. Rates are averages, so if your writing is the kind detectors struggle with, your own chance is higher than the headline. And the fewer students who actually used AI, the larger the share of flags that land on honest work.

This is why Turnitin says its score should not be the sole basis for action, GPTZero says results should not be used to punish or as the final verdict, and OpenAI said its classifier should not be a primary decision-making tool.

Whose writing gets flagged more

The sources agree on the pattern even where their numbers differ.

Second-language writers. In Liang's study, the TOEFL essays that all seven detectors flagged used more predictable word choices. When the authors had ChatGPT enrich the vocabulary, the average false positive rate fell from 61.22% to 11.77%. When they simplified the vocabulary of the eighth-grade essays, it rose from 5.19% to 56.65%. Turnitin's own October 2023 test found no statistically significant gap between second-language and native writers for texts of 300 words or more, and a larger gap below that.

Plain, formulaic or concise writing. OpenAI's help center says its detector research pointed to this kind of writing being hit harder, and that one detector it trained labeled Shakespeare and the Declaration of Independence as AI-generated. Turnitin lists text with little structural variation, text that repeats itself and paraphrase that adds no new ideas.

Short texts. OpenAI called its classifier very unreliable below 1,000 characters. Turnitin needs at least 300 words of prose and says short documents get close to all-or-nothing predictions.

Machine-translated text. Weber-Wulff's team found accuracy on human writing dropped by about 20% once it had been machine-translated into English.

Autistic writers. A 2025 study of about 60,000 Reddit posts found one older open detector flagged under 2% of posts in each group, but significantly more from likely-autistic authors. One detector and one kind of text, so read it as an early signal.

What to do with a number someone shows you

If a teacher, editor or client quotes a detector result, ask four things: which tool, on what date, what the number means in that tool, and what your institution's policy says about scores. Turnitin's percentage, for example, is the share of qualifying prose its model thinks was AI-generated, not necessarily the share of the whole paper, and lists or bullet points are left out.

Then bring your process: version history, notes, an outline and earlier drafts. The page on what to do when a detector flags your essay covers that conversation, and the guide on proving you wrote your essay covers version history in Google Docs and Word.

Don't answer a flag by running your work through rewriting tools. It says nothing about how you wrote it, it can change your meaning, and Turnitin says its English detector also looks for text reworded by AI paraphrasers. If you use an editor on your own draft, HidenGPT included, keep the original beside the edited version, because no editing tool can tell you what a detector will report.

Common mistakes

  • Quoting one headline rate for every detector: name the tool, the date and the kind of text it was tested on.
  • Treating a vendor's figure as independent: check who ran the test and whether the method is published.
  • Reading a Turnitin percentage as the chance you used AI: it is the share of qualifying prose the model flags, and scores from 1% to 19% appear only as an asterisk.
  • Testing one paragraph and trusting the answer: Turnitin needs 300 words, and OpenAI found its classifier very unreliable under 1,000 characters.
  • Rewording a flagged paper before anyone has seen the original: keep the flagged version and its history, and show those first.

Questions

What is Turnitin's AI detection false positive rate?

Turnitin says it is under 1% at the document level, about one in every 100 fully human documents, for documents with more than 20% AI writing detected. It tests each update on more than 700,000 papers written before ChatGPT, shows scores from 1% to 19% only as an asterisk, and says the score should not be the sole basis for action.

Do AI detectors flag non-native English speakers more?

In the best-known study, yes: in 2023 seven detectors flagged 61% of 91 TOEFL essays on average, against about 5% of US eighth-grade essays. Turnitin, which was not part of that study, says its own 2023 test found no significant gap for texts of 300 words or more.

Do AI detectors flag autistic or neurodivergent writers more often?

There is early evidence for autistic writers, but no reliable rate yet. A 2025 study found an older open detector flagged under 2% of Reddit posts overall and significantly more from likely-autistic authors. We did not find a published rate for other neurodivergent groups.

What is Copyleaks' false positive rate?

Copyleaks' detector page claims a .03% false positive rate and over 99% accuracy. That is the company's own figure, and the page we read does not show the test set behind it, so treat it as a vendor claim rather than a measured rate for your writing.

Why did OpenAI shut down its AI text classifier?

Because of its low rate of accuracy, OpenAI said on 20 July 2023. In OpenAI's own tests it caught 26% of AI-written text and labeled 9% of human text as AI-written, and OpenAI's help center now says detectors have not proved reliable in its experience.

Can Grammarly make my essay get flagged as AI?

Grammar and spelling fixes usually don't, according to Turnitin's tests on human-written documents. Turnitin says text produced with Grammarly's generative features, such as drafting, paraphrasing or summarizing, will likely be flagged as AI-generated.

What counts as a false positive in AI detection?

A false positive is fully human-written text that a detector labels as AI-generated. The reverse, AI text passed as human, is a false negative, and Weber-Wulff's 2023 study estimated about 20% of AI-written texts would likely be misattributed to humans, so detectors fail in both directions.

Sources

  1. Liang, Yuksekgonul, Mao, Wu and Zou: GPT detectors are biased against non-native English writers (arXiv 2304.02819, published in Patterns, 2023) (checked October 10, 2026)
  2. Weber-Wulff et al.: Testing of detection tools for AI-generated text, International Journal for Educational Integrity (2023) (checked October 10, 2026)
  3. Chambers and Kelley: The Misclassification of Autistic Writing as AI-Generated (AIED 2025) (checked October 10, 2026)
  4. OpenAI: New AI classifier for indicating AI-written text (31 January 2023, withdrawal note 20 July 2023) (checked October 10, 2026)
  5. OpenAI Help Center: How can educators respond to students presenting AI-generated content as their own? (checked October 10, 2026)
  6. Turnitin Guides: Turnitin's AI writing detection capabilities FAQs (checked October 10, 2026)
  7. Turnitin blog: Understanding the false positive rate for sentences of our AI writing detection capability (14 June 2023) (checked October 10, 2026)
  8. Turnitin blog: New research: Turnitin's AI detector shows no statistically significant bias against English Language Learners (26 October 2023) (checked October 10, 2026)
  9. Copyleaks AI Detector page (checked October 10, 2026)
  10. GPTZero homepage (accuracy FAQ) (checked October 10, 2026)