AI-text detectors flagged more than half of a set of essays written by non-native English speakers
A peer-reviewed 2023 study you can send to a donor, an editor or a university when your own writing is accused of being machine-written. Plus the four things to do before the accusation arrives.
Stanford University
United States
20234 minutes
- Researcher or analyst
- Grants and operations
- Research
- Grants and admin
- Drafting reports and advocacy material
What was measured
On 15 March 2023 the team ran two sets of unambiguously human writing through seven widely used AI-text detectors: 91 TOEFL essays by non-native English speakers, and 88 essays by American eighth-graders.
The detectors were near-perfect on the American children’s essays. On the TOEFL essays they were wrong more than half the time, an average false-positive rate of 61.22%. All seven unanimously flagged 18 of the 91; 89 of the 91 were flagged by at least one.
Then they tested why. Asking a language model to enrich the vocabulary of the TOEFL essays cut the unanimous false flags to a single essay. Asking it to simplify the American children’s essays “as if written by a non-native speaker” pushed misclassification from 5.19% to 56.65%. The detectors are not measuring authorship. They measure linguistic simplicity, and second-language writing is systematically simpler in exactly the way they penalise.
The authors’ conclusion:
First, we strongly caution against the use of GPT detectors in evaluative or educational settings, particularly when assessing the work of non-native English speakers. Our study’s identified high false-positive rate for non-native English writing underscores the potential for unwarranted consequences and the exacerbation of existing biases against these individuals.
An Iraqi researcher, journalist or grant-writer working in a second or third language is more likely to be falsely accused of using AI than a native speaker who actually did. The accusation is close to unfalsifiable: no artefact proves you wrote something yourself. And one prompt bypasses the bias, so a detector cannot catch a deliberate cheat and reliably catches only the honest second-language writer.
Four things to do, none of them technical
- Pre-empt, do not argue afterwards. Before a grant application, article or report leaves your office, paste it into two or three of the free detectors yourself. If your own human-written text scores high, you find out before a programme officer does.
- Keep provenance. Draft in a program that keeps version history and leave it on. A document with weeks of incremental edits is the only evidence that answers “did you write this.”
- Have the citation ready to send back. One paragraph, the free full text of the paper, and the verbatim sentence quoted above. A peer-reviewed line in a Cell Press journal is not something a programme officer can wave away. It is the highest-value item on this page.
- Push for policy, not for an appeal. Ask donors, universities and partners to state in writing that detector output alone will not be treated as evidence of misconduct. The bypass finding is the lever: a tool that cannot catch a cheat and does catch the honest writer is not evidence of anything.
The study nobody has run
The authors released everything under an MIT licence, including the raw essays. Run the identical protocol on Arabic-first or Kurdish-first English writing: around ninety real samples, pasted into the free detectors, tallied in a spreadsheet. An afternoon of clerical work, and it would produce Arabic-language evidence that does not currently exist anywhere. If you do it, publish it.
About the tool named here
ChatGPTChatGPTNever for sensitive materialHas a free version: No payment card needed to sign up. Unlimited basic text chats on the free model, but daily limits on file uploads, image generation, voice chat, and access to the more advanced reasoning models.View tool is an instrument of the experiment here, not the thing being criticised: the researchers used it to rewrite the essays. The seven detectors tested were Originality.AI, Quil.org, Sapling, the OpenAI GPT-2 Output Detector, Crossplag, GPTZero and ZeroGPT. None is in our catalogue, and we do not recommend any of them.
What we do not claim
Everything above is limited by what its sources actually prove. This is the part they do not.
Do not say, in any tense, that "AI detectors flag non-native writers over 60% of the time" as a current general fact. The 61.22% is an average across seven specific tools as they existed on 15 March 2023, measured on 91 TOEFL essays taken from one Chinese educational forum, with no Arabic or Kurdish writers at all, and pulled upward by at least one detector that was never built for text from this generation of models.
Do not claim the bias is unchanged today, and do not claim it has been fixed. The one independent peer-reviewed follow-up we opened (Pratama, PeerJ Computer Science, June 2025) found the direction persists but the size is far smaller and now depends on the tool: a majority-false-accusation rate of 5.56% for non-native versus 2.78% for native authors, a statistically significant bias in one tool and none in two others. That study also found one detector falsely flagging native authors more often than non-native ones. Per-tool generalisation is unsafe in both directions.
Do not claim detector vendors have disputed these findings. We did not verify that.
Do not state that OpenAI withdrew its own classifier for poor accuracy. It is widely reported and probably true, but we could not open a source for it, so it does not appear here as a fact.
This is not a case study of an organisation. Nobody used a tool cleverly here; a university measured a harm that this audience is exposed to.