Are AI Detectors Accurate? What the Evidence Actually Shows in 2026
Written with AI assistance and reviewed by the NorwegianSpark SA editorial team.
Last updated: July 2026
Short answer: no — not accurately enough to be used as proof of anything. The strongest single piece of evidence is not a study by a critic. It is OpenAI's own notice on its own detector: "As of July 20, 2023, the AI classifier is no longer available due to its low rate of accuracy." The company that built ChatGPT built a tool to detect ChatGPT, measured it, and withdrew it. On its own published evaluation, that classifier "correctly identifies 26% of AI-written text (true positives) as 'likely AI-written,' while incorrectly labeling human-written text as AI-written 9% of the time (false positives)."
Everything else in this article is context for that number. Below are the figures from the primary sources — the papers, the vendors' own documentation, and the institutions that tested detectors on real student work — with what each one does and does not establish.
The Peer-Reviewed Picture: Every Tool Below 80%
The most cited independent evaluation is Weber-Wulff and colleagues, published in the International Journal for Educational Integrity in 2023. The team tested 14 detection tools, including the AI-writing features inside two systems already installed across higher education. Their conclusion is stated flatly in the discussion section: "Detection tools for AI-generated text do fail, they are neither accurate nor reliable (all scored below 80% of accuracy and only 5 over 70%)."
Two findings inside that paper matter more than the headline. First, the errors are asymmetric. The tools have, in the authors' words, "a main bias towards classifying the output as human-written rather than detecting AI-generated text" — meaning the typical failure is missing AI text, not flagging innocent writing. That sounds reassuring until you read the second finding: every one of the fourteen tools tested generated false positives, and the risk rose sharply for machine-translated text.
The paper also summarises earlier testing by van Oijen, which found overall accuracy on AI-generated text of just 27.9%, a best-in-class tool topping out at 50%, and roughly 83% accuracy on human-written text — leading to the assessment that the tools were "no better than random classifiers."
So the picture is not "detectors are roughly right with some noise." It is: they miss most AI text, and they still occasionally accuse people who wrote every word themselves.
The Finding That Should Have Ended the Debate
The single most important result in this literature is about who gets falsely accused. In 2023 a Stanford-led team ran seven widely used GPT detectors over two sets of genuinely human-written essays: US eighth-grade essays, and TOEFL essays written by non-native English speakers.
The detectors were near-perfect on the American schoolchildren. On the TOEFL essays they collapsed. The paper reports an average false positive rate of 61.22% — the tools flagged more than half of the human-written non-native essays as machine-generated. Worse, 89 of the 91 TOEFL essays (97.80%) were flagged as AI-generated by at least one detector, and 18 of 91 (19.78%) were flagged by all seven detectors unanimously.
Unanimity is the part that should worry anyone who has ever thought "I'll just check it against a second tool." Running a second detector does not give you independent confirmation. These systems key on the same signal — low lexical variability and predictable word choice — which is exactly what fluent-but-non-native writing looks like. The researchers proved the mechanism by asking ChatGPT to "Enhance the word choices to sound more like that of a native speaker" and re-running the tests: the misclassification rate dropped substantially. The detectors were never measuring authorship. They were measuring vocabulary range.
You do not have to take that from academics. OpenAI reached the same conclusion internally about its unreleased text-watermarking method, and published it: "our research suggests the text watermarking method has the potential to disproportionately impact some groups. For example, it could stigmatize use of AI as a useful writing tool for non-native English speakers." That is the model-maker, in its own words, naming the same bias in a completely different detection technique.
Read the Vendors' Own Numbers Very Carefully
Detection vendors do publish accuracy figures, and the figures are not dishonest. They are just narrower than the way people quote them.
Turnitin, whose AI-writing indicator is the one most students actually encounter, states on its own blog: "Our document false positive rate — incorrectly identifying fully human-written text as AI-generated within a document — is less than 1% for documents with 20% or more AI writing."
Read the qualifier twice. The sub-1% figure is scoped to documents that already contain at least 20% AI writing. It is not a claim about the false-positive rate on a paper a student wrote entirely themselves. In the same post, Turnitin gives a separate figure for the highlighting students actually see: "Our sentence-level false positive rate is around 4%. This means that there is a 4% likelihood that a specific sentence highlighted as AI-written might be human-written."
And Turnitin's own guidance on what to do with the result is more cautious than most of the people citing it: "use the information to initiate a conversation, not to draw a conclusion." The vendor is explicitly telling institutions that its output is not evidence. That instruction is routinely ignored.
What a 1% Error Rate Means at Institutional Scale
Vanderbilt University did the arithmetic in public when it disabled Turnitin's AI detector in August 2023. Its reasoning is worth quoting because it is the clearest illustration of why a small percentage is not a small problem:
"At the time of launch, Turnitin claimed that its detection tool had a 1% false positive rate. To put that into context, Vanderbilt submitted 75,000 papers to Turnitin in 2022. If this AI detection tool was available then, around 750 student papers could have been incorrectly labeled as having some of it written by AI."
Seven hundred and fifty accusations, at one university, in one year, from a rate the vendor considered a selling point. That is the structural problem with detection: the base rate of honest work is enormous, so even an excellent false-positive rate produces a large absolute number of wrongly accused people. And the cost is not symmetric. A missed instance of AI use costs an institution very little. A false accusation costs one specific person a great deal.
Why Detectors Get Fooled So Easily
The other half of the accuracy question is robustness — how the tools behave when someone is actively trying to defeat them. Here the evidence is unambiguous.
The RAID benchmark, released in 2024, is the largest shared evaluation of machine-generated-text detectors: over 6 million generations spanning 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies. The researchers evaluated 8 open-source and 4 closed-source detectors. Their finding: "current detectors are easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models." The paper opens by noting that many commercial and open-source detectors claim accuracy of 99% or more, and that very few of those claims have ever been tested on a shared benchmark.
OpenAI described the same fragility in its own watermarking research, listing how trivially the signal can be stripped: "using translation systems, rewording with another generative model, or asking the model to insert a special character in between every word and then deleting that character."
This is why the "AI humanizer" market exists, and why it works. The detection layer is defeated by anyone who spends thirty seconds trying. It is only reliably triggered by people who were not trying at all — which, again, includes people who simply write plainly.
What the Percentage on the Screen Actually Means
A detector does not identify authorship. It cannot: the text contains no record of where it came from. What these systems output is a statistical estimate of how predictable a passage is — how closely each word follows the word a language model would most likely have chosen. Predictable prose scores as machine-like. That is the whole mechanism.
Several kinds of entirely human writing are highly predictable by construction: technical documentation, legal boilerplate, exam answers written to a rubric, translated text, and — as the TOEFL result showed — competent writing by someone working in a second language. Meanwhile, genuinely machine-generated text becomes unpredictable the moment it is edited, paraphrased or run through another model.
So the number is real, but it is not the number people think it is. "94% AI" does not mean a 94% probability that a machine wrote it. It means the passage sits in a statistical band the tool associates with model output — a band that human writing also occupies.
What to Do Instead
None of this means AI use in written work is a non-issue. It means detection is the wrong instrument, and there are better ones.
- Process beats forensics. Drafts, revision history, outlines and notes are evidence of authorship in a way a probability score is not. Version history in a document is checkable; a detector percentage is not reproducible between tools.
- Ask about the work, not the text. A short conversation about the argument, the sources and the choices made distinguishes an author from a submitter more reliably than any classifier — which is essentially what Turnitin's own "initiate a conversation, not draw a conclusion" guidance amounts to.
- Set an explicit policy and state it up front. Most disputes come from unstated expectations rather than deception. Whether you permit AI for outlining but not drafting, or require disclosure, saying so removes the ambiguity a detector cannot resolve.
- For publishers, judge the output. If you commission writing, the questions that matter are whether the claims are sourced, whether the piece is original, and whether it is any good. Those are answerable. "Did a model touch this?" mostly is not — and for what search engines actually reward, that distinction is covered in our piece on AI content versus human content.
- If you are accused and did the work, the evidence above is public and citable: OpenAI's withdrawal notice, the sub-80% peer-reviewed accuracy finding, the 61.22% false-positive rate on non-native English writing, and the vendor's own statement that its output is a conversation-starter rather than a conclusion. This is general information, not legal or academic advice, and any formal process should be handled through your institution's own appeals procedure.
Where We Might Be Wrong
Two honest caveats, because a piece arguing against overconfidence should not be overconfident itself.
First, the strongest evidence here dates from 2023–2024. Detection research has continued, and it is entirely possible that a tool tested in 2026 performs better than the ones in these studies. What has not changed is the structural problem: any detector that is sensitive enough to catch edited AI text will also flag predictable human text, and the trade-off between those two errors is a property of the task, not of any particular product. We have not found a published, independent 2026 evaluation that overturns the 2023–2024 findings; if one exists, it belongs in this article.
Second, the asymmetry in Weber-Wulff's results cuts both ways. If the dominant error is classifying AI text as human, then a detector flagging a document is comparatively rare — which is precisely why a flag feels so convincing when it happens. That intuition is the trap. Rarity is not reliability, and the Vanderbilt arithmetic shows why.
The Bottom Line
AI detectors are not accurate enough to accuse anyone. The company that makes the most-used AI writer withdrew its own detector for low accuracy. Independent testing put every tool below 80%. The best-documented false-positive pattern falls hardest on people writing in a second language. The leading vendor's own instruction is to start a conversation, not draw a conclusion. And a benchmark of six million samples found the whole category is easily fooled by anyone making the slightest effort.
Used as a rough signal, alongside process evidence and a real discussion, a detector is one weak input among several. Used as proof, it is a coin-flip with a percentage sign on it.
For choosing tools on evidence rather than marketing, see our framework for choosing the right AI tool, our comparison of the best AI writing tools and the major AI chatbots, the practical guidance in AI tools for students and AI tools for marketers, and our guide to learning to code with AI tools. More in the Writing section.
Disclosure: NeuralPuls holds no affiliate relationship with any AI-detection tool, AI writing tool or AI model provider. Nothing in this article is a paid placement, and there is no commission attached to any position it takes.