AI Tools11 min read

Do AI Detectors Actually Work? What the Evidence Shows in 2026

AI text detectors are being used to fail students and reject job applicants. The research on whether they can do the job is not encouraging, and the errors fall hardest on people writing in a second language. Here is how they work and what their scores really mean.

A student is called into an academic integrity meeting. The evidence is a percentage from a detection tool. They wrote the essay themselves, over four evenings, and now have to prove a negative against a number nobody in the room can explain.

This is happening at scale, and the tools driving it are considerably weaker than their marketing suggests. Here is what they actually measure, what the published evaluations found, and what any of it justifies.

What a Detector Is Measuring

No detector reads text and knows where it came from. They estimate, and almost all of them estimate from two related statistical properties.

Perplexity is a measure of how surprising each word is given the words before it. Language models are trained to predict likely continuations, so their output tends to sit in the high-probability lane: fluent, expected, low surprise. Human writing wanders more. It reaches for the odd word, breaks its own rhythm, includes the phrasing that a probability model would not have picked. Low perplexity is read as a machine signal.

Burstiness is the variation in sentence length and structure across a passage. Human prose is uneven. A long, clause-heavy sentence gets followed by a short one. Model output has historically been flatter, with sentences clustering around a similar length and shape. Low burstiness is read as a machine signal too.

Some tools add a trained classifier on top, fitted on labelled samples of human and AI text. Some use a second model to score how likely the first model would have been to produce this exact string.

Every one of these approaches shares a structural problem: they detect the *style* of a probable-word generator, not the *origin* of the text. Those are different things, and the gap between them is where all the damage happens. A careful, plain, evenly-paced human writer produces exactly the statistical fingerprint these tools were built to flag.

OpenAI Shut Down Its Own Detector

The clearest signal about the state of this technology came from the company with the most to gain from it working.

In January 2023 OpenAI released an AI Text Classifier. Its own published evaluation was blunt about the numbers: on a challenge set, the tool correctly identified about 26 percent of AI-written text as likely AI-generated, and incorrectly labelled about 9 percent of human-written text as AI-generated.

Read those together. It missed roughly three quarters of the machine text it was pointed at, and still managed to falsely accuse close to one in eleven human writers.

In July 2023 OpenAI withdrew the tool, saying it was being discontinued due to its low rate of accuracy. That is a company with full access to the models being detected, extensive training data on both sides of the problem, and an obvious commercial interest in shipping a working detector, concluding it could not build one good enough to keep online.

Commercial detectors that stayed on the market publish accuracy figures well above that. Those figures are usually generated in-house, on datasets the vendor selected, using unedited model output. Independent evaluations tend to land lower, and the drop is steepest once the AI text has been lightly edited by a human, which is how virtually all of it reaches the real world.

The Bias Finding That Should Have Ended the Practice

In July 2023, Weixin Liang and colleagues at Stanford published a study in the journal Patterns titled "GPT detectors are biased against non-native English writers".

They ran seven widely used GPT detectors over two sets of essays. One set was TOEFL essays written by non-native English speakers. The other was essays by US eighth-graders, all native speakers.

The native-speaker essays were classified almost perfectly.

More than half of the TOEFL essays were flagged as AI-generated. Across the detectors, all seven unanimously misclassified a substantial share of them.

The mechanism is exactly the one described above. Second-language writing typically uses a more common vocabulary and steadier sentence construction, because that is what fluency looks like on the way up. Those are the two properties every perplexity-and-burstiness detector treats as machine evidence.

The researchers then demonstrated the second half of the problem. When they asked a language model to rewrite the flagged TOEFL essays with more literary vocabulary, the detection rate collapsed. The tools were fooled by the trivial countermeasure, while continuing to flag the honest writers.

That combination is the worst possible one. The false positives fall on people least equipped to contest them, and the actual evasion is one prompt away.

What a Percentage Score Does Not Mean

The interface is doing a lot of damage on its own. A detector that outputs "87% AI" is presenting an estimate with the visual authority of a measurement.

Things that number is not:

  • A probability that the text is AI-generated. It is a classifier output on an internal scale, not a calibrated posterior.
  • A proportion of the text written by AI. An 87 percent score does not mean 87 percent of the words came from a model.
  • Evidence with a chain of custody. There is no artefact behind it. Nothing was found. A model produced a number.
  • Reproducible. Run the same passage through three detectors and you can get three materially different answers. Run it through the same detector after a tool update and the answer can change again.

A number with none of those properties should not, on its own, decide an academic integrity case or a hiring outcome. Several universities reached the same conclusion and turned the feature off. Vanderbilt University publicly disabled Turnitin's AI detection in 2023 citing false positive concerns, and it was not alone.

If You Have Been Falsely Flagged

Being accused of something you did not do, with a number as the evidence, is genuinely difficult to argue against. Some things help.

Lead with process, not with the score. Arguing about detector accuracy sounds like a defence. Showing your work is a demonstration. Version history in a document editor that tracks revisions, saved drafts with timestamps, outlines, research notes, browser history from the relevant days: any of these show writing happening over time, which no detector output can contradict.

Ask what the threshold is and where it came from. Most people applying these tools cannot answer, and the question moves the conversation from your character to their evidence. Ask what false positive rate the institution accepts and how it was established.

Cite the Stanford finding if it applies to you. If you write in English as a second language, the published research says the tool is more likely to flag you regardless of what you wrote. That is a documented property of the instrument, not a claim about your case.

Offer to demonstrate. A supervised writing sample on a related topic, or simply talking through the argument and sources in the piece, tends to settle it fast. Someone who wrote something can discuss why they structured it that way. That conversation is far better evidence than any score in either direction.

Keep drafts going forward. Unfair as it is, the practical protection is a documented process, and it only works if it exists before the accusation.

The Honest Position on Using AI to Write

The detection question and the ethics question get tangled together, and they are separate.

Detectors do not work well. That is a fact about the tools. It is not a licence to submit generated text as your own where that is prohibited, and it does not make undisclosed use fine because you are unlikely to be caught.

The workable line most institutions and publications have converged on:

  • Generally accepted: brainstorming angles, outlining structure, explaining a concept you are trying to understand, checking grammar, tightening a draft you wrote, generating example material clearly labelled as such
  • Generally prohibited: submitting model output as your own work, generating citations without verifying they exist, using it where an explicit policy forbids it
  • Almost always right: disclosing when you are unsure, and reading the actual policy rather than assuming it

The verification point deserves emphasis. Language models fabricate citations confidently and fluently, and a fabricated reference in submitted work is a problem entirely independent of any detector. Check every source you did not personally read.

Using AI as a Tool Without Handing Over the Writing

The useful mental model is that a language model is good at the parts of writing that surround the writing. Getting unstuck on an opening. Listing the objections to your argument so you can answer them. Explaining a concept three ways until one lands. Rewriting a paragraph you already wrote to be shorter.

Generai is an iPhone app for that kind of work: chat for thinking through and drafting, plus image generation from text prompts. Ask it to argue the opposite side of your thesis, or to list what a sceptical reader would push back on, and you get something that improves work that stays yours.

Two limits worth stating plainly. It will state incorrect things fluently, so anything factual needs checking against a real source. And it cannot tell you what your institution's policy is. That part is on you to read.

The Short Version

  • Detectors measure statistical style, not origin, and those are not the same thing
  • OpenAI built one, published a 26 percent detection rate with 9 percent false positives, and withdrew it as too inaccurate
  • A Stanford study found seven detectors flagged more than half of non-native English TOEFL essays as AI, while classifying native-speaker essays near-perfectly
  • Light editing defeats detection, so the tools catch honest writers more reliably than dishonest ones
  • A percentage score is an uncalibrated estimate with no evidence behind it, not a measurement
  • Keep draft history. It is the only defence that works, and it has to exist before you need it
  • None of this makes undisclosed use acceptable. Read the actual policy and verify every citation

Frequently Asked Questions

Are AI detectors accurate?

Not reliably enough to act on a single result. The strongest signal about their limits came from OpenAI, which built its own AI Text Classifier and then withdrew it in July 2023, stating in its own announcement that it was being discontinued because of its low rate of accuracy. In its published evaluation the tool correctly flagged only about 26 percent of AI-written text while incorrectly flagging about 9 percent of human writing. Independent evaluations of commercial detectors report wide variation depending on the model, the topic and how much the text was edited.

Do AI detectors discriminate against non-native English speakers?

The evidence says yes. A 2023 study led by Weixin Liang at Stanford, published in Patterns, ran seven widely used GPT detectors over TOEFL essays written by non-native English speakers and over essays by US eighth-graders. More than half the non-native essays were misclassified as AI-generated, while the native-speaker essays were classified near-perfectly. The likely cause is that these detectors reward unusual word choice and varied sentence structure, and second-language writing tends to be more measured in both.

Can I prove I wrote something myself?

Process evidence is the strongest thing available, and it has to exist before the accusation. Draft history in a document editor with version tracking, notes, outlines, browser research history and dated backups all show work happening over time. A detector score is a probability estimate with no supporting evidence attached, so a documented drafting trail is generally more persuasive than arguing with the number.

Is using AI to help write something actually cheating?

That depends entirely on the rules of the place you are submitting to, and those rules vary enormously. Many universities and publications now permit AI assistance for brainstorming, outlining or editing while prohibiting submitted text generated wholesale, and most ask for disclosure. The practical advice is to read the specific policy rather than assume, and to disclose when unsure, because undisclosed use is what turns a permitted tool into a violation.

Try Generai: AI Chat & Art Creator

Mentioned in this article. Download free from the App Store.

More Articles