Skip to content
Sound More Human

Blog

For teachers: talking to a student about a suspected AI paper

· 6 min read

By

A paper does not sound like the student. The prose is polished, the argument is generic, and a detector has returned a high score. What happens next is a conversation, and how you run it decides whether you learn what happened or only damage a working relationship.

We make a rewriting tool for students, so we are not a neutral party, and this article is about your process rather than our product. It rests on published research and on guidance from universities and from a detector vendor. Where it depends on a source, the source is linked.

What a detector score can and cannot show

A score is a statistical statement about a piece of text. It is not a statement about who wrote it, how it was produced or what the writer intended. The research on how reliable the statement is points one way.

  • The vendor says it is not a verdict. In a March 2023 post, Turnitin wrote that it does not make a determination of misconduct (opens in a new tab), acknowledged that its false positive rate is not zero, and told instructors to apply professional judgment, knowledge of their students and the context of the assignment. It also claimed a false positive rate below 1 per cent. That is the vendor’s own figure, and independent testing is what to weigh it against.
  • Independent tests find serious errors in both directions. Weber-Wulff and colleagues tested 14 tools, including Turnitin and PlagiarismCheck, and concluded they are neither accurate nor reliable (opens in a new tab). In the paper, six of the fourteen produced false positives on human-written or machine-translated text. The tools were also more likely to call text human than to catch AI-written text, so a low score is no clearance either. The authors’ conclusion is that these tools should not be used in academic settings.
  • The errors are not evenly spread. A Stanford study tested seven detectors on 91 TOEFL essays by non-native English speakers and 88 essays by US eighth-graders. The detectors were near-accurate on the eighth-graders. On the TOEFL essays the average false-positive rate was 61.3 per cent, and 97.8 per cent of the essays were flagged by at least one detector, though all were human-written. Liang et al., Patterns, 2023 (opens in a new tab). The samples were small, and the authors say so, but the direction of the error is the point.
  • Some institutions acted on this. Vanderbilt University disabled Turnitin’s AI detector (opens in a new tab) in August 2023, citing among other reasons that the vendor gives no detailed account of how it decides. UMass Amherst’s Undergraduate Student Success Center advises instructors not to rely solely on AI detection tools (opens in a new tab) and to talk to the student before any formal report.

A fair summary: a score can be a reason to look closer. It cannot show that a student used a tool, and it cannot show that a student did not. For a longer treatment, Are AI detectors accurate? covers the research.

Before you speak to the student

  • Name your actual concern. “The score is high” is not a concern. “The argument in section two does not match what she said in seminar” is. Write down what you observed, separately from the number.
  • Check what you promised. Reread your syllabus and assignment sheet. If the rules on AI were unclear, a fair conversation starts there.
  • Gather your own comparison material. Earlier work by the same student, in-class writing, discussion posts. Vanderbilt’s guidance to its instructors suggests comparing against the student’s previous style, tone and level.
  • Know your institution’s procedure and the point at which you are required to report. Do not improvise the formal steps.
  • Decide in advance what would change your mind. If you cannot name evidence that would clear the student, you are not investigating.

What to ask

Turnitin’s own advice for these conversations is to assume positive intent and to be open that false positives happen, because otherwise the meeting turns defensive. Ask questions that a student who wrote the paper can answer and one who did not may find hard, without announcing which you expect.

  1. Can you walk me through how you approached this assignment, from the first idea?
  2. Where did you find these sources, and what did you take from this one?
  3. What is the strongest objection to your argument, and how would you answer it?
  4. Here is a sentence from page three. What did you mean by it, in other words?
  5. Do you have notes, an outline or earlier drafts? Could I see the version history?
  6. Did you use any tools, including grammar, translation or AI tools? What did each do?

Read the answers for fluency about the ideas, not for polish. A student who can explain the argument and its weak points has given you real evidence. A student who cannot has given you a reason to keep asking. It is still not proof, because nerves, a language barrier or a paper written weeks ago can all produce a blank.

A fair process

  • Hear the student before deciding anything, and give them time to bring drafts and notes rather than requiring answers on the spot.
  • Do not act on a score alone. Every source above says the score needs support, and the vendor says so too.
  • Put the evidence in writing: what you observed, what the student said, what they produced.
  • Say what the outcome options are and how the student can appeal before you meet, not after.
  • Treat “unresolved” as an allowed result. Sometimes the evidence does not settle it. Saying so is better than a finding you cannot defend.

Second-language writers

The Stanford result is why extra care is warranted here: plain, careful prose from a writer working in a second language is what the tested detectors flagged most. The Weber-Wulff team saw a related effect, with accuracy on human writing falling by about 20 per cent once it had been machine-translated into English, a common route for students drafting in their first language. Reading a second-language writer’s clean prose as suspicious is a bias with a documented mechanism, so a flag on such a paper needs more corroboration, not less. Our article on second-language writing and detector flags covers what students can do and how they might explain themselves.

Making the next assignment easier to trust

  • Say in the syllabus what is allowed, per assignment. Vanderbilt tells its instructors to communicate early and set clear guidelines. A rule nobody wrote down is hard to enforce fairly.
  • Assess process, not only product. Weber-Wulff and colleagues end by recommending that written assessment focus on how a student’s skills develop rather than on the final paper alone. Outlines, annotated sources, dated drafts and a short reflection all make a paper easier to stand behind.
  • Tie tasks to your classroom. Vanderbilt suggests in-class writing and topics discussed in class, which a generic essay cannot answer.
  • Ask for a short oral check on a sample of papers, chosen by rule and not by suspicion, so no student is singled out.

If you are deciding what to say about tools like ours, we publish a page for educators at /educators that sets out what a rewriting tool can and cannot show. The short version is that a rewrite proves nothing about who wrote a text, and we do not claim otherwise. Students who want to do the right thing can read our guide to reading a course’s AI policy, which tells them to check your syllabus and to ask you when it is silent.