Research · Facts and data sheets

What the research says about AI writing, and about detectors.

Peer-reviewed numbers on how machine text differs from human text, how detectors read it, and where they get it wrong. Every figure carries its source.

Cited sources
13
Detector families
4
Last reviewed
Oct 2026

§ 01

Key findings

Twelve numbers worth knowing.

Published results on how AI text differs from human text, and on how well detectors tell them apart. Each card names its source; the full list is at the end of the page.

01Structure

0.0macro-F1

Structure alone, with no word choice at all, separates AI from human B2B blog posts. After the AI rewords its own posts, it still scores 96.1.

Original posts97.0
After the AI rewords its own posts96.1

Macro-F1, 0–100. Rewording barely moves it.

[1] SlopShape v3, arXiv 2609.15369, Sept 2026

02Readers

0.0%

Readers who use AI heavily caught AI articles 92.7% of the time, with 4.0% false positives. Everyone else was at chance: 56.7% caught, 51.7% false positives.

Chance

Heavy AI users

Everyone else

0%100%
AI articles caught Human articles wrongly flagged

[3] Russell, Karpinska & Iyyer, ACL 2025

03Bias

0%

of essays by non-native English writers were flagged as AI by perplexity-based detectors.

[4] Liang et al., 2023

04Paraphrase

0.0%0.0%

DetectGPT's detection rate at 1% false positives, before and after one paraphrase pass.

Before70.3%
After one paraphrase pass4.6%

[5] Krishna et al., NeurIPS 2023

05Watermarks

0.0%0.0%

Detection of a text watermark (SynthID) before and after one paraphrase pass.

Before66.5%
After one paraphrase pass1.5%

[6] DAMAGE authors' test

06Polishing

43–0%

of human-written texts were flagged by leading detectors after “extremely minor” AI polishing.

43%65%
0%human texts flagged100%

[7] Saha & Feizi, APT-Eval 2025

07Fine-tuning

0%0%

A leading detector's flag rate, before and after per-author fine-tuning.

Before97%
After per-author fine-tuning3%

[11] Chakrabarty et al., 2025

08Grammar

0.0×

GPT-4o uses trailing “-ing” analysis clauses (“…, highlighting the need for…”) at 5.3 times the human rate.

[8] Reinhart et al., PNAS 2025

09Vocabulary

0×

“Delves” appeared at 28 times its expected rate in 2024 PubMed abstracts.

[9] Kobak et al., Science Advances 2025

10Platforms

~0%

fewer views for posts flagged with LinkedIn's “seems like AI slop” report, according to LinkedIn. The report launched 30 Jul 2026.

[10] LinkedIn, via TechCrunch

11Regulation

2 Aug 2026

EU AI Act Article 50 transparency duties for AI-generated content start to apply.

[12] EU AI Act, Article 50

12Fidelity

“14 days” → “14 weeks”

A humanizer we tested quietly changed a duration in its output. Rewording is cheap; changing a fact is expensive. That's why every rewrite here is checked against its source.

Source

Returns are accepted within 14 days of delivery.

Humanizer output

We accept returns within 14 weeks of delivery.

Fact changed · a fact check would block this

Example sentence. The changed value is from our test

[13] Our lab test, Oct 2026

§ 02

How AI detectors work

Four ways to guess who wrote it.

Every detector on the market leans on one or more of these families. Each measures something real, and each has a documented way to fail.

Fig. 2.1 · Statistical

Schematic, not to scale

What it measures

How predictable each word is to a language model. When nearly every next word is the obvious one, the text reads as machine-made.

We treat predictability as one weak signal, never as a verdict on its own.

Where it holds and where it breaks

  • 61%[4]

    of essays by non-native English writers flagged as AI. Careful, plain prose is predictable too.

  • 70.3% → 4.6%[5]

    DetectGPT's detection rate at 1% false positives after one paraphrase pass.

§ 03

AI vs human

The measured differences.

The clearest gap isn't vocabulary. It's how the argument is put together, and whether a reader is left with anything to do.

Table 1 · Structure

Share of B2B posts showing each pattern[2]

AI posts Human posts
  • No way for the reader to act

    2.6×

    97%
    38%
  • Thesis announced up front

    1.8×

    93%
    51%
  • Summary section

    3.3×

    88%
    27%
  • Close restates the thesis

    6.4×

    77%
    12%
  • Old-versus-new framing

    2.9×

    76%
    26%

[2] Sitefire / SlopShape

Table 2a · Grammar

Trailing “-ing” analysis clauses

0.0×

GPT-4o, relative to the human rate.

Revenue rose 12% in Q3, highlighting the strength of the new strategy.Example sentence, for illustration

[8] Reinhart et al., PNAS 2025

Table 2b · Vocabulary

“Delves”

0×

2024 PubMed abstracts, relative to the expected rate.

This review delves into the mechanisms underlying…Example sentence, for illustration

[9] Kobak et al., Science Advances 2025

Table 3 · Rhythm, range and specifics

What a reader feels before they can name it

Illustration, not measured data

Sentence-length spread

Default model prose keeps sentences in a narrow band. People swing from four words to forty.

AIHuman040 words

Emotional range

Default model prose stays level and upbeat. People get annoyed, wry, unsure, then certain.

start of textend

Specifics

Default model prose generalises. People name the client, the date, the number, the street.

AI

  • a client
  • recently
  • significant growth
  • many teams

Human

  • the Leeds office
  • 3 March
  • £41,200
  • invoice #2291

§ 04

How we benchmark

Measure the misses, not just the hits.

A detector that is right 95% of the time can still be wrong about most of the people it flags. Our method is built around false positives.

  1. Step 1

    Held-out human corpus

    Writing we never train on: native and non-native authors, lightly edited drafts, technical and casual registers. It's where false positives hide.

  2. Step 2

    Calibrate to the writer

    Score new text against the same writer's known samples, not against an imagined average human. A plain writer is allowed to stay plain.

  3. Step 3

    Fidelity checks

    Every rewrite is diffed against its source for numbers, units, names and links. Any change fails the call instead of shipping.

  4. Step 4

    Report false positives

    False-positive rates per group at fixed thresholds, published next to detection rates. Never a single accuracy number.

Fig. 4 · Detect benchmark

What our benchmark report will show

In progress, not yet published
  1. 1Corpus make-up

    Which human and AI texts are in the held-out set, and how they were kept out of training.

  2. 2False positives by group

    Native and non-native writers, lightly edited drafts and technical writing, each at fixed thresholds.

  3. 3Generic vs calibrated

    The same texts scored with and without the writer's own baseline, side by side.

  4. 4One paraphrase pass

    How much detection drops after a single rewording, the test word-level detectors fail.

  5. 5Fidelity error rate

    How often a rewrite changes a number, unit, name or link, and how often the check catches it.

We'll publish the numbers here, good or bad, with the corpus description and thresholds. Until then, the figures on this page are from the cited studies.

§ 05

Data sheets

Take it to your next meeting.

Short PDFs that summarise each topic with full citations. In production now.

Coming soon

DS-01 · PDF

Detector families and where they fail

Statistical, classifier, watermark and structure detectors: what each measures and the cited failure rates.

Coming soon

DS-02 · PDF

Structure signals in B2B writing

The five patterns that separate AI and human posts, with rates and a reviewer's checklist.

Coming soon

DS-03 · PDF

Fidelity: what rewriting tools change

Why numbers, units, names and links must survive a rewrite, and how a fact check fails closed.

Coming soon

DS-04 · PDF

Detect benchmark method

Held-out corpus, writer calibration and per-group false-positive reporting. Draft.

§ 06

Responsible use

Built for writers. Not for cheating.

We sell to publishers, content teams and people writing in their own voice. Some uses are off the table, whatever plan you're on.

What IntactVoice is for

  • Publishers and editors checking drafts before they ship
  • SEO and content teams keeping several writers' voices distinct
  • Founders and executives writing in their own voice, faster
  • Agencies writing for a client who has agreed to a voice profile
  • Product teams adding review signals to their own tools

What it is not for

  • Students, schools and universities, or any coursework, exam or assessment
  • Government bodies, or getting around any government or licensing requirement
  • Imitating a real person without their consent
  • Treating a Detect score as proof of who wrote something
  • Disciplining students or staff on a score alone
  • Hiding that content was machine-made where the law requires disclosure

Signals, not verdicts

Detect returns likelihoods and the reasons behind them, sentence by sentence. A person makes the call.

Consent for voices

A voice profile is built from a writer's own samples, with their agreement. Profiles can be deleted at any time.

Facts locked

Rewrite will not change a number, unit, name or link. If it would, the call fails and shows you the diff.

Use outside these lines breaks our terms and ends access. Questions about a use case? Write to support@intactvoice.com.

§ 07

References

Figures are quoted as published. Where a source reports a range or a platform's own claim, we say so.

  1. [1]

    SlopShape v3, arXiv 2609.15369, Sept 2026

    Structure-only classifier on B2B blog posts; macro-F1 before and after the model rewords its own posts.

  2. [2]

    Sitefire / SlopShape

    Share of AI and human posts showing each structural pattern.

  3. [3]

    Russell, Karpinska & Iyyer, ACL 2025

    Human readers as detectors; heavy AI users versus everyone else.

  4. [4]

    Liang et al., 2023

    Perplexity-based detectors on essays by non-native English writers.

  5. [5]

    Krishna et al., NeurIPS 2023

    Effect of one paraphrase pass on DetectGPT at a 1% false-positive rate.

  6. [6]

    DAMAGE authors' test

    Effect of one paraphrase pass on a text watermark (SynthID).

  7. [7]

    Saha & Feizi, APT-Eval 2025

    Human texts with “extremely minor” AI polishing, scored by leading detectors.

  8. [8]

    Reinhart et al., PNAS 2025

    Grammatical and rhetorical features of model text versus human text.

  9. [9]

    Kobak et al., Science Advances 2025

    Excess vocabulary in 2024 PubMed abstracts.

  10. [10]

    LinkedIn, via TechCrunch

    “Seems like AI slop” report, launched 30 Jul 2026; reach figure as stated by LinkedIn.

  11. [11]

    Chakrabarty et al., 2025

    Per-author fine-tuning and a leading detector's flag rate.

  12. [12]

    EU AI Act, Article 50

    Transparency obligations for AI-generated content apply from 2 Aug 2026.

  13. [13]

    Our lab test, Oct 2026

    A third-party humanizer changed a duration in its output. Internal test, not peer reviewed.

Detect API

See the reasons, not just the score.

Run Detect on your own writing and get sentence-level signals with the pattern behind each one.