Why AI detectors disagree, and what they actually measure

Every AI detector measures a stand-in for authorship, and each one measures a different stand-in. That’s why two tools can read the same paragraph in opposite ways, and why a score is a signal you weigh, not proof.

IntactVoice research team

Sources checked Oct 4, 2026

Published
Updated
Reading time
7 min
On this page
  1. Five ways to measure, five different questions
  2. Why short text is unreliable
  3. False positives land on predictable writers
  4. What the best results show
  5. How to read a score
  6. Limits of this guide
  7. Build your own baseline
  8. Questions people ask
  9. Sources

On October 4, 2026 we sent the same 197-word sample to two AI detectors. One, a paid detection API, returned a score we read as about 88% AI. The other, a free public website, reported 0% AI and 100% human-written. Same text, same day, opposite answers.

Neither result tells you who wrote the sample. (It was drafted inside an AI-assisted project, so neither tool had a clean case.) What the pair shows is that a detector score depends on which detector you ask, and at 197 words both were working near the edge of what text that short can support.

Figure

One 197-word sample, two detectors

Same text, same day: October 4, 2026

Detector A

Paid detection API

≈ 88% AI

HumanAI

Detector B

Free public website

0% AI

HumanAI
Source: our test. Detector A's scale is inferred from its website; neither tool is named because one sample says nothing about either tool's accuracy.
Show as table

Five ways to measure, five different questions

No detector identifies an author. Each one measures something that tends to go with machine writing.

MethodWhat it measuresExamplesWhere it fails
Statistical (“zero-shot”)How predictable each word is to a language modelDetectGPT, Binoculars, early GPTZeroOne paraphrase breaks it; flags plain, formulaic and non-native writing
Trained classifiersHow closely the text matches the default assistant voice, over several sentencesPangram, GPTZero, Turnitin, Originality.ai, CopyleaksUnreliable on short text; often flags human text lightly polished by AI; misses 10% to 18% when a model imitates a writer’s style
WatermarksA pattern the model provider hid in its own word choices, readable only with its keyGoogle’s SynthID; Anthropic announced one for Claude in August 2026Only covers that provider’s models; needs a reasonable length; a full rewrite removes it; few people can run the check
StructureHow the argument is organised: announced thesis, summary stage, restated closeSlopShape (research)Unreviewed; called about 7% of human posts AI in its own test
Expert readersVocabulary clusters, formulaic structure, missing originalityEditors and heavy AI usersSlow, and only as good as the reader
Based on the published research summarised in our research page. Statistical methods are now mostly found in free tools; commercial vendors use trained classifiers.

The first two families respond to fluent, low-surprise prose sustained over many sentences. They don’t look for a word list or for em dashes. That’s why one edited sentence rarely changes a verdict, and why an AI rewrite of a whole human paragraph often does.

Why short text is unreliable

Detectors need enough text to separate a pattern from chance, and the vendors say so in their own documentation:

Figure

How much text detectors ask for

Published minimum or test length, in words (approximate where the source counts characters or tokens)

Pangram minimum50 words

50

Copyleaks minimum350 characters, about 60 words

60

Watermark paper's test length200 tokens, about 150 words

150

Our two-detector sample197 words

197

Turnitin minimum300 words of prose

300
Sources: Pangram 3.2 model card; Copyleaks Help Center; Kirchenbauer et al. 2023; Turnitin AI Writing Report guide. Read Oct 4, 2026.
Show as table
  • Pangram’s model card sets a minimum input of 50 words.
  • Turnitin asks for “at least 300 words of prose text in a long-form writing format”.
  • Copyleaks requires at least 350 characters, about 60 words, and suggests testing around 350 words.
  • Watermark checks need length too. The original watermark paper (Kirchenbauer et al.) ran its main tests on 200-token passages, roughly 150 words.

An early study found that both people and automatic detectors get better as excerpts get longer, and that even multi-sentence excerpts fooled expert raters over 30% of the time (Ippolito et al., ACL 2020). A short social post or a product description can fall below most of these floors.

False positives land on predictable writers

The best-known result is Liang et al. (2023): seven detectors flagged essays by non-native English writers as AI at an average rate of 61.22%. That bias belonged mainly to perplexity-based tools, which mistake plain vocabulary for machine text. Retrained commercial classifiers report much lower rates on these groups, though those figures are mostly the vendors’ own.

Other patterns hold across studies:

  • Formulaic and concise writing is flagged more, because it’s predictable. Constrained genres such as recipes, legal boilerplate and product specs are hard cases.
  • Light AI polish on human writing gets flagged. In APT-Eval (Saha and Feizi, 2025), “extremely minor” polishing by the small Llama-2-7B model got 32.31% of human texts flagged by ZeroGPT, 42.56% by Pangram and 64.71% by GPTZero, all early-2025 versions.
  • Base rates multiply small errors. When Vanderbilt disabled Turnitin’s AI detector in August 2023, it noted that a 1% false-positive rate across its 75,000 yearly submissions would mean about 750 papers wrongly flagged.

What the best results show

The strongest current classifiers are much better than the 2023 tools. OpenAI withdrew its own classifier in July 2023 after it caught only 26% of AI text. In an Epoch AI test reported in July 2026, Pangram, GPTZero and Originality.ai missed at most 0.7% of plainly prompted AI text, and Pangram and GPTZero flagged none of 495 human passages.

The same test shows the limit. When models were given five samples of an author’s writing and asked to imitate it, Pangram missed 10%, GPTZero 11% and Originality.ai 18%; in scientific writing the misses rose to 24% to 29%. Going further, fine-tuning a model on an author’s complete works cut detection from 97% to 3% in a 2025 study by Chakrabarty and colleagues. And readers who use AI heavily themselves did better than most tools: a majority vote of five such readers misclassified 1 of 300 articles in Russell, Karpinska and Iyyer (ACL 2025).

Detection of plain, unedited assistant output is close to solved. Detection of writing as people actually produce it, edited, mixed and imitated, isn’t.

How to read a score

Checklist6 checks

  • Check the length first. Below the tool’s own minimum, set the score aside.
  • Note the tool and its version. Scores from different tools aren’t on the same scale, and a model update can move them.
  • Look for reasons. A tool that highlights sentences and says why gives you something to check; a single percentage doesn’t.
  • Run a second tool. If two detectors disagree, treat the question as open.
  • Compare against the writer’s own baseline: how does the same tool score pieces you know they wrote?
  • Never act on a score alone. Turnitin says its score “should not be used as the sole basis for adverse actions”, and GPTZero suggests using its result “as a conversation starter, and not as the final verdict”.

Limits of this guide

  • Our two-detector example is one 197-word sample on one day. We’ve left both tools unnamed because one result was read from an API score whose scale we had to infer, and a single sample says nothing about either tool’s general accuracy.
  • Most vendor accuracy figures are measured by the vendors on data they chose. Independent tests of the newest production models are scarce.
  • Detectors and models both change every few months. Dates matter on every number here.

Build your own baseline

Pick five pieces a person on your team wrote before 2022, each over 300 words, and run them through the detector you use. Write down the scores. That’s your false-positive baseline for that tool on your kind of writing, and it’s the number to remember the next time a score looks alarming.

Questions people ask

Can a detector prove a text was written by AI?

No. Every method gives a probability or a signal. Watermarks come closest, because they test for a specific provider’s key, but they only cover that provider’s models and disappear after a full rewrite.

Why did my own writing get flagged?

The usual reasons are length (too short to judge), predictability (plain, formulaic or template-driven text) and AI polishing (a grammar or “improve” pass that rewrote sentences). Run a longer sample, and check whether any tool rewrote the text before you scanned it.

Do platforms run these detectors on posts?

Some do. Substack has let readers run a Pangram scan on posts longer than 100 words since July 21, 2026 (Mashable), and LinkedIn runs its own classifiers alongside a report button added on July 30, 2026. Google has published no statement that it uses AI detectors for search ranking; see what Google says about AI content.

Sources

Dates are each source's publication or last-updated date, or the day we read it.

  1. 1
    Our test, October 4, 2026

    one 197-word synthetic sample sent to a paid detection API and a free public detector website. Internal, not peer reviewed

  2. 2
  3. 3
  4. 4
    Pangram 3.2 model card

    Pangram · Feb 27, 2026

  5. 5
    Using the AI Writing Report

    Turnitin · read Oct 4, 2026

  6. 6
    Minimum character count for the AI Detector

    Copyleaks Help Center · read Oct 4, 2026

  7. 7
    A watermark for large language models

    Kirchenbauer et al., ICML · 2023

  8. 8
  9. 9
    GPTZero’s AI detection technology

    GPTZero · read Oct 4, 2026

  10. 10
  11. 11
    New AI classifier for indicating AI-written text

    OpenAI · Jan 31, 2023

    Withdrawn Jul 20, 2023

  12. 12
    AI text detectors struggle when language models mimic an author’s style

    The Decoder · Jul 19, 2026

    Reporting an Epoch AI study

  13. 13
  14. 14
  15. 15
    How Claude’s text watermark works

    Anthropic · Aug 14, 2026

  16. 16

Rewrite with your facts locked.

IntactVoice moves a draft toward a writer's own habits and checks every number, unit, name and link against the source. Plans are billed by the word.