In the spring of 2023, William Quarterman opened his student portal at the University of California, Davis and found a failing grade attached to a cheating accusation. His take-home history exam had been run through an AI detector called GPTZero, and the software said he had used a chatbot. He hadn’t. Quarterman was referred to the university’s judicial affairs office, and it took weeks before the case collapsed. Weeks later, a classmate named Louise Stivers went through the same fog after Turnitin flagged part of her Supreme Court case brief. She cleared her name only by handing over her Google Docs revision history.
Those cases are three years old, and the tools have changed. The problem they exposed has not. Schools now run student writing through AI detection software as a matter of routine, and most families have never been told how those scores are produced, how often they are wrong, or what happens when one lands on a child who did nothing.
What these tools actually measure – and what they don’t
An AI detector does not keep a database of machine-generated essays and check your child’s work against it. That is how plagiarism software works, and it is why Turnitin’s plagiarism checker can point to the exact source a student copied. AI detection is a different animal. It scores text on statistical properties, most commonly perplexity (how predictable each next word is) and burstiness (how much sentence length and rhythm vary from line to line).
Fluent, even, predictable prose reads as machine-like. That turns out to describe a lot of writing by humans. When Communications of the ACM surveyed the field in June 2026, it noted a list of works these tools have flagged as AI: Shakespeare’s Macbeth, the U.S. Constitution, and the Book of Genesis.
So the score is not a finding of fact. It is a probability, tuned by a threshold its makers can change, and it says “this reads like AI” far more reliably than it says “this was written by AI.”

The accuracy numbers are worse than the marketing suggests
Vendors market near-perfect accuracy. Independent testing keeps returning something else. The most useful studies for parents are the ones that report both error directions – how often a detector flags human writing, and how often it misses actual AI text.
| Study (date) | What was tested | What it found |
|---|---|---|
| International Journal for Educational Integrity (Feb 2026) | Turnitin and Originality on 192 mixed human, AI, and hybrid texts | Originality reached 0.69 accuracy, Turnitin 0.61. Both were near-useless on hybrid text – partly AI, partly human |
| University of Florida, IEEE Symposium on Security and Privacy (2026) | Five commercial detectors on scientific writing | False-positive rates ranged from 0.05% to 68.6%; false negatives from 0.3% to 99.6%. A trivial wording change wrecked them |
| University of Pennsylvania RAID benchmark (2024–25) | 12 detectors against 6 million+ AI samples | Simple edits – synonyms, reordering, lookalike characters – pushed error rates above 95% |
| Advances in Physiology Education (2024) | Four detectors vs. human raters on 190 student essays | Individual detectors had roughly 1.3% false positives; combining three or more drove the joint false-positive rate toward zero |
Figures are as reported by each study on the dates shown. Percentages are not directly comparable across rows because the test sets differ.
Read the first two rows together and a pattern appears. The tools are decent at catching raw, untouched chatbot paste. They fall apart the moment text has been edited, translated, or blended with human writing – which is exactly what ordinary student drafting looks like.
Who gets caught in a false positive
Errors are not spread evenly. They cluster around students whose writing style trips the same statistical wires as a language model.
Non-native English speakers are the best-documented group. A 2023 Stanford-led study found early GPT detectors misclassified more than half of essays written by non-native writers. The 2026 study in the International Journal for Educational Integrity found a borderline-significant trend toward the same bias, and its authors warned that “over-reliance on machine-based identification risks misclassifying legitimate student work.”
Neurodivergent students show up in the case files too. At the University of New South Wales in 2024, a student named Rushi Vyas was flagged after submitting a carefully drafted assignment; he told reporters he believed the tool was stacked against him. He was cleared after walking reviewers through his draft history.
Then there is the stranger finding. A 2026 paper analyzing 135,389 pairs of manuscripts from a professional editing service found that simply polishing a human-written text – the kind of clean-up a writing center does – could push its AI score up in some detectors and down in others. Editing style, not authorship, was moving the needle.

How common is this in schools now?
Very. A RAND survey released September 30, 2025 found that 54% of students and 53% of English, math, and science teachers had used AI for school – jumps of 15 points or more in a year or two. Only about a third of teachers reported having a school or district policy on academic integrity and AI. RAND also found students naming, unprompted, that they worry about being falsely accused.
Adoption of detectors specifically ran even harder. Center for Democracy & Technology data puts regular use of AI detection tools among K–12 teachers at 68% in the 2023–24 school year, up from 38% the year before. Discipline followed: 64% of teachers said a student at their school had been disciplined over AI use.
Policy is starting to catch up, unevenly. New York City Public Schools announced a moratorium on student-facing generative AI in grades 2K–8 for the 2026–27 school year, with limited, approved use in high school. California’s Department of Education released a model AI policy on June 25, 2026 that explicitly addresses limits on AI detection software and parent review rights. Wyoming’s draft framework goes further, requiring written parental consent before AI tools are used in instruction.
The case for keeping the tools – carefully
It would be sloppy to end this as a simple “detectors are junk.” The honest research picture is more interesting, and a parent who ignores it will lose an argument they could have won.
The June 2026 study that followed 1,163 master’s theses found false positives were “almost absent” for several modern tools, a genuine improvement over the 2023-era results. Turnitin maintains its false-positive rate is under 1% for academic writing over 300 words. And the physiology-education study showed that when you require three detectors to agree, the joint false-positive rate collapses toward zero while real AI text still gets caught.
That is the useful framing: a detector can be a smoke alarm. It cannot be a judge. Nick Diakopoulos, a professor at Northwestern University, told Communications of the ACM that the tools “are not reliable enough to make strong assertions or to sanction people.” His recommendation – and the recommendation running through nearly every serious study – is agreement across multiple tools, followed by a human conversation and evidence of the writing process.

What students can do
- Keep your process. The clearest thing that saved falsely accused students was not an argument. It was evidence: Google Docs version history, earlier drafts, outlines, notes, browser history. Write in a document that timestamps your revisions rather than pasting a final draft into a blank file.
- Check your own work before you submit. If you want to see roughly what your teacher’s dashboard shows, you can run your draft through an AI detector free of charge. Treat the number as a diagnostic, not a verdict – a high score does not mean you cheated, and a low one does not mean you’re safe.
- Know what’s allowed. “Don’t use AI” and “use AI but disclose it” are very different rules. Get the policy in writing before the assignment, not after.
- Ask for a human review, not a re-score. If you are accused, the winning move is to request that a person examine your drafts and answer questions about your own work.
As for the question every student eventually asks – can I just use AI and humanize the text to beat the detector? Technically, yes. Researchers have shown that paraphrasing and rewriting drop detection to near zero. It also means your writing is no longer yours, and it does nothing for the thing school is for. It is a way to win a game nobody is grading you on.
What parents can ask

Most districts have no script for this yet, so your questions will do the work. The useful ones are specific:
- Which detection tool is used, and is it the only evidence, or one factor among several?
- What was the actual score, and what threshold caused the flag?
- Can my child see the highlighted text and respond to it in a conversation?
- What process evidence can we provide – drafts, revision history, notes?
- Which district policy governs this, and can I read it?
If the answer to the first question is “the tool alone,” you are on solid ground. Turnitin itself says its AI report “should not be used as the sole basis for adverse actions against a student.” In the 2026 Adelphi University case, a judge found the school’s cheating allegations against a student named Orion Newby “completely false” and ordered his record expunged – after the school relied on a single Turnitin result and ignored two other checks that said the paper was human-written.
Questions parents and students keep asking
Can an AI detector prove a student cheated?
How accurate are these tools, really?
Why does the software flag work my child wrote by themselves?
What should we do if a child is falsely accused?
Are schools even allowed to use these detectors?
What actually protects students
The detector arms race is a dead end for schools, because it asks the wrong question. The evidence is clear that the tools misread real human writing, that they punish students who polish their work or write in a second language, and that determined cheaters can slip past them in seconds. What the tools reward is compliance with a statistical pattern, not learning.
The schools handling this well are moving in a different direction: assignments that build in drafting, reflection, and in-class work; policies that spell out what AI use is allowed and how to disclose it; and a hard rule that no single score can end a student’s semester. Detection can flag a conversation. It cannot replace one.