AI models have stopped using em dashes, but you can still spot their writing. A new study from marketing firm Graphite found 13,000 phrases that show up at least twice as often in AI-generated text as in human writing, according to TechCrunch AI. Most of the classic giveaways are gone. Each new model brings its own habits to replace them.
The study also shows the major labs moving in different directions. “It turns out that Claude models are actually getting closer to the human word distribution over time,” Graphite’s chief AI officer Greg Druck told TechCrunch. “And for the GPT models, it’s getting further away.”
🔬 How Graphite Ran the Test
Comparing AI and human writing at scale is harder than it sounds. If you just hand a model a topic, you can’t tell whether a word choice comes from the model or from the subject. Graphite set up the study to deal with that:
- Human baseline: The team collected 10,000 articles published before ChatGPT launched, so they’re almost certainly human-written.
- AI rewrites: Several frontier models rewrote each article from a summary instead of from the original text, which kept the source wording from leaking through.
- Comparison: With matched human and AI versions of the same pieces, the team measured how often specific words, phrases and sentence structures appeared in each.
Graphite counted a phrase as a “tell” if it appeared at least twice as often in AI output as in human text.
📊 Each Model’s Habits
Claude Opus 5.5 has three notable tells:
- “this matters” appears 116x more often in Claude’s output than in human writing
- “why X matters” appears 92x more often
- “dependable” appears 23x more often
Opus 5.5 has mostly dropped the old “it’s not X, it’s Y” pattern. It now leans on a close relative instead: something “is more than an X, it’s a Y.” Its strongest habit is telling readers why a topic is important, over and over.
OpenAI’s Astra has different habits. It likes to describe “another dimension” of a topic and hedges with phrases like “may provide” or “can provide.” Its signature tells like “not simply X” and “rather than relying on X” appear 100x+ more often than in human writing. Its biggest tell is what Graphite calls “corrective framing,” where the model defines something by what it isn’t before saying what it is.
✂️ Em Dashes Are Nearly Gone
Every lab seems to have heard the complaints about em dashes:
- Opus 5.5 uses them 99% less than Opus 5 did.
- Astra uses them 88% less often than the human samples.
- Gemini 3.1 Pro has dropped them almost entirely.
That sounds like progress, but Druck says the overall count of tells isn’t going down. “They are managing to remove the most well-known tells, but other ones pop up. And every model version has its own,” he told TechCrunch.
To me, that’s the most important result in the study. The labs are fixing the tells people complain about, not the underlying tendency that produces them. Remove one habit and the model finds another.
🧩 Why the Labs Can’t Fully Fix It
Anthropic and OpenAI both say their newest models write more naturally. Anthropic says Opus 5.5 “communicates more naturally than prior models.” OpenAI promised “fewer odd turns of phrase” with the GPT-6 versions of Sol and Luna.
Druck doesn’t think either lab can get there completely. “A general hypothesis I have is that the labs are less able to control some of these things than you might expect,” he said. “These are giant models with billions of parameters. They have some finite number of tests they can run, and things slip through.”
🛠️ What You Can Do With This
If you publish AI-assisted content, the findings point to a few practical steps:
- Edit for patterns, not just words. Banning “delve” won’t do much. Look for repeated sentence structures, especially contrast framing (“not X but Y,” “more than X”).
- Know your model’s habits. Claude’s tells differ from GPT’s. Build a short checklist of phrases for each model you use.
- Recheck after model updates. A new version means new quirks. Last quarter’s checklist will slowly go out of date.
- Treat AI detectors with caution. If the tells change with every release, detectors trained on older models will miss newer ones.
⚠️ Limitations
The study comes from a marketing firm, and TechCrunch’s coverage doesn’t say whether it was peer-reviewed. Rewriting articles from summaries is a sensible control, but it’s still one specific task. Models may write differently in emails, code comments or casual chat. The 2x cutoff for a “tell” is also a fairly low threshold, which helps explain how the list reached 13,000 phrases.
If the pattern holds, removing AI tells won’t be a one-time job. Editors will need to keep updating what they look for with each model release. TechCrunch AI’s original report has more detail on Graphite’s findings.