Evidence, not vendor claims

Are AI detectors accurate?

Short answer: accurate enough to be useful as a prompt for a conversation, nowhere near accurate enough to be used as proof. The best peer-reviewed evidence shows false-positive rates that make single-score decisions indefensible — especially for non-native English writers.

The headline number

61% of human essays flagged as AI

In 2023 a Stanford team ran 91 human-written TOEFL essays through seven widely used AI detectors. The detectors classified them as AI-generated at an average rate of 61.22%. The same detectors, given essays by native English-speaking US eighth-graders, misclassified only 5.19%.

Bar chart comparing average false-positive rates across seven AI detectors: 61.22% for non-native English writers' TOEFL essays versus 5.19% for native English writers' US eighth-grade essays.

Two further results from the same study matter as much as the headline:

That last pair is the whole problem in one experiment. The detectors are not finding machine authorship. They are finding plain, predictable language — and then labelling it machine authorship.

What the tools measure

Perplexity and burstiness, not authorship

Three cards explaining that AI detectors measure perplexity (how surprising each word is) and burstiness (how much sentence rhythm varies), and that clear, well-edited human writing scores low on both.

No detector has access to a document's authorship. It has access to the text. Everything it reports is inferred from statistical properties of that text — chiefly how predictable each word is given the ones before it, and how much sentence length and rhythm vary.

Human writing that is clear, well-edited, formulaic by genre, or written by someone with a smaller working vocabulary in English scores as "predictable" on exactly these measures. So does AI writing. The measure cannot separate them, and no amount of model improvement changes what is being measured.

What a 1% error rate means at scale

Why institutions have switched detectors off

Turnitin has stated a roughly 1% false-positive rate for its AI detector. Vanderbilt University worked out what that meant for them: they submitted about 75,000 papers to Turnitin in 2022, so a 1% rate implies roughly 750 papers wrongly flagged in a single year at a single institution. In August 2023 Vanderbilt disabled the detector, citing that arithmetic, the lack of any published explanation of how the detector works, and the documented bias against non-native English writers.

Their conclusion was blunt: they did not believe AI detection software was an effective tool that should be used.

This is the number that gets lost in vendor marketing. A 99%-accurate test still produces hundreds of false accusations at institutional scale, and each false positive lands on one specific person who has to prove a negative.

Reading a score honestly

What a detector result can and cannot support

Can support

A reason to look more closely. A prompt to ask about process, drafts and sources. A signal worth pairing with other evidence.

Cannot support

A finding of misconduct on its own. A grade penalty. A hiring decision. Any claim about who typed the words.

Never assume

That two detectors agreeing means more than one. The Stanford data shows they fail together on the same texts.

Last reviewed: 14 August 2026. We update these pages when detector vendors change their claims or new peer-reviewed evidence is published.

FAQ

Common questions

What is the most reliable AI detector?

No detector has demonstrated accuracy sufficient to serve as proof of authorship. Independent testing published in Patterns (Cell Press) in 2023 found an average false-positive rate of 61.22% on human-written TOEFL essays across seven widely used detectors. Treat any vendor accuracy claim as unverified unless it is backed by independent, peer-reviewed testing on writing similar to yours.

Does running a second detector confirm the first?

No. In the Stanford study all seven detectors unanimously misclassified 18 of 91 human-written essays. Detectors share similar underlying measures, so they tend to fail on the same texts. Agreement between them is not independent confirmation.

Why do AI detectors flag non-native English speakers so often?

Because they measure how predictable text is, not who wrote it. Writing with simpler vocabulary and more uniform sentence rhythm scores as machine-like. When researchers made non-native essays more linguistically complex, the false-positive rate dropped from 61.22% to 11.77% without changing the author.

Has any university stopped using AI detection?

Yes. Vanderbilt University disabled Turnitin's AI detector in August 2023, citing the absence of a published explanation of how it works, documented bias against non-native English writers, and the fact that Turnitin's own stated 1% false-positive rate would have meant roughly 750 wrongly flagged papers out of the 75,000 the university submitted in 2022.