Skip to content

Blog

Are AI detectors accurate? What the research says

· 8 min read

A detector returns a number, and a number looks like a measurement. It is worth knowing what that number is actually built on before anyone acts on it — particularly when the action is an academic misconduct case.

What a detector is measuring

Detectors do not recognise text that a model produced. They score how predictable a passage is: whether each word is the one a language model would expect, and how much the sentence lengths vary. Predictable, even prose scores as machine-written. Surprising, uneven prose scores as human.

That is a proxy, and the gap between the proxy and the question people ask it — who wrote this? — is where every problem below comes from. Plain, careful writing is exactly what the method flags, whoever produced it.

What the research found

The most cited independent test is by Weixin Liang and colleagues at Stanford, published in Patterns in 2023. They ran seven publicly available detectors over TOEFL essays written by non-native English speakers, and over essays by US eighth-graders.

  • On the TOEFL essays — human-written, every one — the detectors averaged a false-positive rate above 60 per cent. More than nine in ten of those essays were flagged by at least one detector.
  • On the native-speaker essays, false positives were close to zero.
  • The authors’ explanation is the mechanism above: a writer working in a second language tends to use commoner words and steadier sentence shapes, which is precisely what these tools score as machine-like.

Liang et al., “GPT detectors are biased against non-native English writers” is open access, and worth reading in full if you are deciding policy. The Markup followed the same finding into real cases, where international students were accused on the strength of a detector score.

The same paper reports the opposite failure too: lightly rephrasing genuinely machine-written text pushed it under the detectors’ thresholds. So the tools were over-flagging one group of human writers while missing the thing they exist to find.

If you are a teacher or an administrator

The honest reading is that a detector score is a prompt to look closer, never a finding. Some universities reached that conclusion and switched their detection features off rather than defend decisions they could not explain.

  • Never open a case on a score alone. Ask what else supports it: draft history, the student’s other work, a conversation about the argument they made.
  • Know which students the errors fall on. A policy driven by these scores will land hardest on second-language writers and on anyone whose natural register is plain.
  • Say in the syllabus what is allowed and how it is checked. Most disputes come from a rule nobody wrote down.
  • Assess process where you can — outlines, drafts, an oral defence of the thesis. It is harder to fake and easier to defend.

If you are a student who has been flagged

Being flagged is not evidence of anything, and you can say so with a citation. Keep your draft history — version history in a document editor is the single most useful thing you can produce — and ask what the accusation rests on besides the score. If you write in English as a second language, the Stanford finding is directly about you.

Where we stand, since we sell a rewriting tool

Sound More Human makes no claim about what any detector will say about a rewrite, and never will. That is not modesty; it is that no such claim can be honest. Detector behaviour changes without notice, the tools disagree with each other, and as the research above shows they are unreliable in both directions. Anyone promising you a particular score is selling something they cannot deliver.

What a rewriting tool can do is make prose less flat to read — the habits that cause the flatness are specific and fixable. That is a writing-quality argument, and it is the only one we make.