On this page
On October 4, 2026 we sent the same 197-word sample to two AI detectors. One, a paid detection API, returned a score we read as about 88% AI. The other, a free public website, reported 0% AI and 100% human-written. Same text, same day, opposite answers.
Neither result tells you who wrote the sample. (It was drafted inside an AI-assisted project, so neither tool had a clean case.) What the pair shows is that a detector score depends on which detector you ask, and at 197 words both were working near the edge of what text that short can support.
Figure
One 197-word sample, two detectors
Same text, same day: October 4, 2026
Detector A
Paid detection API
≈ 88% AI
Detector B
Free public website
0% AI
Show as tableHide table
| Detector | Reported | Read as |
|---|---|---|
| A: paid detection API | score 88.02 | about 88% AI |
| B: free public website | 0% AI, 100% human | 0% AI |
Five ways to measure, five different questions
No detector identifies an author. Each one measures something that tends to go with machine writing.
The first two families respond to fluent, low-surprise prose sustained over many sentences. They don’t look for a word list or for em dashes. That’s why one edited sentence rarely changes a verdict, and why an AI rewrite of a whole human paragraph often does.
Why short text is unreliable
Detectors need enough text to separate a pattern from chance, and the vendors say so in their own documentation:
Figure
How much text detectors ask for
Published minimum or test length, in words (approximate where the source counts characters or tokens)
Pangram minimum50 words
Copyleaks minimum350 characters, about 60 words
Watermark paper's test length200 tokens, about 150 words
Our two-detector sample197 words
Turnitin minimum300 words of prose
Show as tableHide table
| Tool or test | Length |
|---|---|
| Pangram minimum | 50 words |
| Copyleaks minimum | 350 characters, about 60 words |
| Watermark paper's test length | 200 tokens, about 150 words |
| Our two-detector sample | 197 words |
| Turnitin minimum | 300 words of prose |
- Pangram’s model card sets a minimum input of 50 words.
- Turnitin asks for “at least 300 words of prose text in a long-form writing format”.
- Copyleaks requires at least 350 characters, about 60 words, and suggests testing around 350 words.
- Watermark checks need length too. The original watermark paper (Kirchenbauer et al.) ran its main tests on 200-token passages, roughly 150 words.
An early study found that both people and automatic detectors get better as excerpts get longer, and that even multi-sentence excerpts fooled expert raters over 30% of the time (Ippolito et al., ACL 2020). A short social post or a product description can fall below most of these floors.
False positives land on predictable writers
The best-known result is Liang et al. (2023): seven detectors flagged essays by non-native English writers as AI at an average rate of 61.22%. That bias belonged mainly to perplexity-based tools, which mistake plain vocabulary for machine text. Retrained commercial classifiers report much lower rates on these groups, though those figures are mostly the vendors’ own.
Other patterns hold across studies:
- Formulaic and concise writing is flagged more, because it’s predictable. Constrained genres such as recipes, legal boilerplate and product specs are hard cases.
- Light AI polish on human writing gets flagged. In APT-Eval (Saha and Feizi, 2025), “extremely minor” polishing by the small Llama-2-7B model got 32.31% of human texts flagged by ZeroGPT, 42.56% by Pangram and 64.71% by GPTZero, all early-2025 versions.
- Base rates multiply small errors. When Vanderbilt disabled Turnitin’s AI detector in August 2023, it noted that a 1% false-positive rate across its 75,000 yearly submissions would mean about 750 papers wrongly flagged.
What the best results show
The strongest current classifiers are much better than the 2023 tools. OpenAI withdrew its own classifier in July 2023 after it caught only 26% of AI text. In an Epoch AI test reported in July 2026, Pangram, GPTZero and Originality.ai missed at most 0.7% of plainly prompted AI text, and Pangram and GPTZero flagged none of 495 human passages.
The same test shows the limit. When models were given five samples of an author’s writing and asked to imitate it, Pangram missed 10%, GPTZero 11% and Originality.ai 18%; in scientific writing the misses rose to 24% to 29%. Going further, fine-tuning a model on an author’s complete works cut detection from 97% to 3% in a 2025 study by Chakrabarty and colleagues. And readers who use AI heavily themselves did better than most tools: a majority vote of five such readers misclassified 1 of 300 articles in Russell, Karpinska and Iyyer (ACL 2025).
Detection of plain, unedited assistant output is close to solved. Detection of writing as people actually produce it, edited, mixed and imitated, isn’t.
How to read a score
Checklist6 checks
- Check the length first. Below the tool’s own minimum, set the score aside.
- Note the tool and its version. Scores from different tools aren’t on the same scale, and a model update can move them.
- Look for reasons. A tool that highlights sentences and says why gives you something to check; a single percentage doesn’t.
- Run a second tool. If two detectors disagree, treat the question as open.
- Compare against the writer’s own baseline: how does the same tool score pieces you know they wrote?
- Never act on a score alone. Turnitin says its score “should not be used as the sole basis for adverse actions”, and GPTZero suggests using its result “as a conversation starter, and not as the final verdict”.
Limits of this guide
- Our two-detector example is one 197-word sample on one day. We’ve left both tools unnamed because one result was read from an API score whose scale we had to infer, and a single sample says nothing about either tool’s general accuracy.
- Most vendor accuracy figures are measured by the vendors on data they chose. Independent tests of the newest production models are scarce.
- Detectors and models both change every few months. Dates matter on every number here.
Build your own baseline
Pick five pieces a person on your team wrote before 2022, each over 300 words, and run them through the detector you use. Write down the scores. That’s your false-positive baseline for that tool on your kind of writing, and it’s the number to remember the next time a score looks alarming.
Questions people ask
Can a detector prove a text was written by AI?
No. Every method gives a probability or a signal. Watermarks come closest, because they test for a specific provider’s key, but they only cover that provider’s models and disappear after a full rewrite.
Why did my own writing get flagged?
The usual reasons are length (too short to judge), predictability (plain, formulaic or template-driven text) and AI polishing (a grammar or “improve” pass that rewrote sentences). Run a longer sample, and check whether any tool rewrote the text before you scanned it.
Do platforms run these detectors on posts?
Some do. Substack has let readers run a Pangram scan on posts longer than 100 words since July 21, 2026 (Mashable), and LinkedIn runs its own classifiers alongside a report button added on July 30, 2026. Google has published no statement that it uses AI detectors for search ranking; see what Google says about AI content.
Sources
Dates are each source's publication or last-updated date, or the day we read it.
- 1Our test, October 4, 2026
one 197-word synthetic sample sent to a paid detection API and a free public detector website. Internal, not peer reviewed
- 2GPT detectors are biased against non-native English writers
Liang et al., Patterns · 2023
- 3Automatic detection of generated text is easiest when humans are fooled
Ippolito et al., ACL · 2020
- 4Pangram 3.2 model card
Pangram · Feb 27, 2026
- 5Using the AI Writing Report
Turnitin · read Oct 4, 2026
- 6Minimum character count for the AI Detector
Copyleaks Help Center · read Oct 4, 2026
- 7A watermark for large language models
Kirchenbauer et al., ICML · 2023
- 8Almost AI, almost human: the challenge of detecting AI-polished writing (APT-Eval)
Saha and Feizi, arXiv · May 2025
- 9GPTZero’s AI detection technology
GPTZero · read Oct 4, 2026
- 10Guidance on AI detection and why we’re disabling Turnitin’s AI detector
Vanderbilt University · Aug 16, 2023
- 11
- 12AI text detectors struggle when language models mimic an author’s style
The Decoder · Jul 19, 2026
Reporting an Epoch AI study
- 13Readers prefer outputs of AI trained on copyrighted books over expert human writers
Chakrabarty et al., arXiv · revised Mar 2026
- 14People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
Russell, Karpinska and Iyyer, ACL · 2025
- 15How Claude’s text watermark works
Anthropic · Aug 14, 2026
- 16Substack adds tool that detects AI-generated content
Mashable · Jul 22, 2026