Why every brand starts to sound the same with AI

When most teams draft with the same few models, the drafts drift toward each other. The studies that measured it, why it happens, and the one input that stays yours.

IntactVoice research team

Sources checked Oct 4, 2026

Published
Updated
Reading time
5 min
On this page
  1. What the studies measured
  2. Why it happens
  3. What doesn’t converge
  4. A working method
  5. Limits of this evidence
  6. Next step: put your outlines side by side
  7. Sources

Ask two content teams to write about the same topic with the same model and you’ll get two drafts with the same outline, the same examples and often the same opening move. Several studies have now measured this, in fiction, in essays and in company blog posts, and they agree on the direction: AI help makes each piece better on average and the whole set more alike.

What the studies measured

The clearest experiment is Doshi and Hauser (Science Advances, July 2024). 293 people wrote eight-sentence stories; some could ask GPT-4 for story ideas. With up to five AI ideas available, outside judges rated the stories 8.1% more novel and 9.0% more useful, with most of the gain going to the less creative writers. The same stories were also more alike. The authors measured how close each story sat to the others in its group: access to AI ideas raised that similarity by an amount equal to 10.7% of the full spread of scores in the human-only group (one AI idea) and 8.9% (five ideas).

Figure

Better stories, closer together

293 writers, eight-sentence stories, with and without GPT-4 story ideas

Each story, rated by judges

Novelty
+8.1%
Usefulness
+9.0%

With up to five AI ideas, against writing alone.

The stories as a group

One AI idea
10.7%
Five AI ideas
8.9%

Rise in similarity to the other stories, as a share of the human-only range.

Source: Doshi and Hauser, Science Advances, Jul 12, 2024. Dot panels are a schematic, not plotted data.

Other studies point the same way:

  • Padmakumar and He (ICLR 2024) found that writing essays with an instruction-tuned model, InstructGPT, made different authors’ essays more alike, while writing with the base GPT-3 model didn’t. The loss of diversity came from the text the model contributed.
  • The Artificial Hivemind study (NeurIPS 2025) put 26,000 open-ended prompts to many models and found both repetition within each model and “strikingly similar outputs” across different models.
  • Graphite’s September 2026 study, vendor research with a single fixed prompt, found that “every model’s writing is closer to at least one other model than to human writing.”
  • In SlopShape (September 2026), human company blog posts sat at the 83.5th rarity percentile on average and AI posts on the same topics at the 43.6th: the human posts were much less like their neighbours.

So switching models doesn’t get you far. The models resemble each other more than any of them resembles a person.

Why it happens

The best-supported explanation is how chat models are trained after pre-training. Tuning a model on human preferences narrows the range of next words it considers likely. One study measured the effective number of options shrinking by a factor of 2 to 5 overall, and from about 12 to 1.2 at the start of a response (Yang, Li and Holtzman, 2025). Raters tend to prefer familiar, organised, agreeable text, and the models learn that preference. Every major model goes through a version of this process, which is why they converge on each other.

For a content team, the practical result is simple. The first idea a model gives you is the first idea it gives your competitors, and asking it to “be more creative” doesn’t change much. Giving it samples of your writing helps less than you’d hope for blog posts: in Wang et al. (2025), models given five samples of an author’s writing got close on news and email style and stayed generic on blogs and forum posts.

What doesn’t converge

The Padmakumar and He result holds the useful detail: the human-written parts of those essays stayed diverse. Sameness enters through the text the model supplies. Material the model never saw can’t converge with anyone else’s, because nobody else has it.

The model can supplyOnly your team can supply
A statistic anyone can citeYour own counts: tickets, refunds, hours, prices, conversion rates
“Many teams struggle with…”The exact question a customer asked last week, in their words
The standard advice on the topicWhere you disagree with it, and what happened when you tried
A generic exampleA named tool, a date, a version number, a mistake you made
A tidy summaryA caveat you’d only know from doing the work
The right-hand column is what the SlopShape study found marks human posts: real dates, named sources, things to use, a challenged default and the reader’s own situation.

This is also what readers and search engines are looking for. Google’s own guidance asks whether a page offers “original information, reporting, research, or analysis” (see our guide to what Google says about AI content).

A working method

  1. Collect material before you open a model. Answer four questions in a few lines each: What are the real numbers? What’s one specific case, with the names of things? What do we do differently from the standard advice? What can the reader use directly: a link, a template, a script, a price?
  2. Ask for several angles and drop the first. Have the model list three or four approaches, name the one everyone would write (usually the listicle or the “ultimate guide”), and pick another that your material supports.
  3. Write the draft from your notes. Keep your sentences wherever they exist. Use the model to order, cut and connect them.
  4. Don’t ask for a polish of the whole piece. A full rewrite replaces your wording with the model’s, which is exactly the part that converges. Ask for specific fixes instead.
  5. Compare with what already ranks. Search your topic, open the top five results and write down their headings. If your outline matches theirs, change the angle before you edit a sentence.

Limits of this evidence

  • Doshi and Hauser studied short stories by online participants with ideas from GPT-4, not company blogs. The effect sizes are modest.
  • Graphite’s model comparison is vendor research on one prompt; SlopShape is an unreviewed single-author preprint from a company that sells a checker.
  • No study we know of has measured brand-voice convergence in marketing directly. We’re applying results from neighbouring settings, and we’d change this guide if better evidence appeared.

Next step: put your outlines side by side

Take your three most recent posts and the closest post from two competitors on each topic. Copy only the headings into a table, one column per post. If the headings could swap between columns without anyone noticing, write your next post from the four intake questions above before you ask a model for anything.

Sources

Dates are each source's publication or last-updated date, or the day we read it.

  1. 1
  2. 2
  3. 3
    Artificial Hivemind: the open-ended homogeneity of language models (and beyond)

    Jiang et al., NeurIPS Datasets and Benchmarks · Oct 27, 2025

  4. 4
    AI Tells

    Graphite · Sep 16, 2026

    Vendor study, one fixed prompt

  5. 5
    SlopShape: identifying AI-generated commercial web content, v3

    arXiv 2609.15369 · Sep 28, 2026

    Jochen Madler (Sitefire)

  6. 6
    LLM probability concentration: how alignment shrinks the generative horizon

    Yang, Li and Holtzman, arXiv · Jun 2025 (revised Sep 2026)

  7. 7
  8. 8
    Creating helpful, reliable, people-first content

    Google Search Central · updated Oct 1, 2026

Rewrite with your facts locked.

IntactVoice moves a draft toward a writer's own habits and checks every number, unit, name and link against the source. Plans are billed by the word.