The AI Tells Nobody Agrees On (and the Ones the Data Actually Supports)
There's a lot of confident folk wisdom floating around about how to spot AI-written text: count the em dashes, look for "delve," check if every sentence has three items in it. Some of that folk wisdom holds up under scrutiny. Some of it doesn't. I went looking for the actual research behind the most common claims, built a tool to test them against real text, and the tool immediately proved one of my own assumptions wrong. Here's what I found, what I built, and what I'd get wrong if I trusted intuition alone.
The short version
AI writing tells fall into four categories with very different levels of evidence behind them:
- Lexical overuse — specific words with measured, statistically significant frequency spikes since ChatGPT's launch. This is the best-documented category.
- Structural and rhetorical patterns — predictable sentence architectures ("it was not X, it was Y," three-item lists) that show up repeatedly in AI-generated prose but haven't been measured with the same rigor as vocabulary.
- Punctuation habits — the em dash, mainly. Genuinely measurable, but weaker evidence than people think, and getting weaker as vendors make the habit suppressible.
- Image artifacts — a moving target. The tells that worked in 2023 (six fingers, garbled text) are mostly fixed in current models. The reliable checks have shifted to physics and provenance.
I turned all four into a free tool: AI Tell Checker. More on that below, including a real test that exposed a gap in my own first version.
What the research actually shows
Vocabulary: the strongest signal
This is the one category where the evidence is genuinely rigorous. A longitudinal study of PubMed abstracts tested 61 words suspected of being AI-favored for statistically significant year-over-year growth after ChatGPT's public launch, and found 10 with significant excess growth compared to control words, including "delve," "commendable," and "intricate." A separate analysis of more than 15 million PubMed abstracts quantified just how extreme some of these spikes got: "delve" ran at roughly 28 times its expected 2024 frequency, "underscore" about 14 times, and "showcasing" about 11 times.
Two caveats worth taking seriously. First, this list isn't static: "delve" usage has already dropped since early 2024, once people started publicly mocking it, while other AI-favored words like "significant" kept climbing. Second, the effect isn't limited to a handful of famous words. Researchers using a formal three-step method to isolate AI-driven frequency spikes in scientific writing identified 21 "focal words" showing the same pattern, not just the two or three that get memed on social media.
Practical takeaway: a word list is the cheapest, best-evidenced place to start a purge, but it needs to be revisited periodically. Treat it as a living list, not a fixed blacklist.
Structure and rhetoric: real, but less measured
Contrast framing ("it was not X, it was Y"), rule-of-three lists, and indiscriminate emphasis all show up constantly in AI-generated prose, and there's a useful crowd-sourced observation worth repeating here: text that emphasizes ordinary, low-stakes statements as if they were important is often a tell, because the model doesn't actually know what matters and hedges by emphasizing everything.
This category is real but harder to measure than vocabulary. Nobody has run a PubMed-scale frequency study on three-item lists. That doesn't make the pattern less common in AI text, it just means the evidence here is closer to "widely observed by editors and writers" than "peer-reviewed and quantified." Worth flagging, worth fixing, but hold it more loosely than the vocabulary findings.
The em dash: real effect, overstated as proof
The em dash is the most talked-about AI tell and also the most contested. There's a real, measurable generational shift behind it: one controlled test generating 10 stories each from GPT-3.5, GPT-4o, and GPT-4.1 found 0 em dashes from GPT-3.5, 14 from GPT-4.1, and 16 from GPT-4o. There's also a plausible mechanical explanation: in GPT-4's tokenizer, a leading space plus em dash is a single token, while a comma-and-conjunction construction usually costs two or three tokens, making the dash a cheap, high-probability choice for a next-token predictor.
But multiple independent sources push back hard on treating dash frequency as proof of anything. Human formal writers, journalists, and 19th-century novelists all use em dashes heavily, and as of late 2025, OpenAI's CEO confirmed that instructing ChatGPT to avoid em dashes via custom instructions now actually works, which means the "tell" is becoming something a user can simply turn off. Density is worth tracking against your own baseline. Treating it as a verdict is not defensible.
Image tells: a checklist with an expiration date
If you're checking images instead of text, the checklist from two years ago is mostly obsolete. By 2026, image generators have solved the obvious problems: six-fingered hands, garbled background text, and waxy over-smooth skin have all mostly been fixed. Skin texture in particular has improved through deliberate grain and pore simulation, and eye reflections that used to be a reliable giveaway now render correctly in most current models.
What still works: hands in complex poses (closed fists, interlaced fingers), asymmetric or malformed ears, background text and signage (still one of the most frequent flaws in 2026), and shadow or reflection physics that don't match a single consistent light source. The newer layer of detection has moved past visual inspection entirely, toward reverse image search for compositing tells and cryptographic provenance standards like SynthID and C2PA.
What happened when I tested my own tool
I built a checker that scores text against seven categories: banned vocabulary, stock AI phrases, em dash density, contrast framing, rule-of-three lists, reflexive hedging, and markdown emphasis markers. First test: I fed it a paragraph I knew was 100% AI-generated, a routine staff email about procurement deadlines.
It came back "Moderate," 6.9 flags per 1,000 words. But when I looked at the breakdown, six of the seven categories returned zero. The only thing it caught was a single em dash.
Two separate bugs were hiding behind that one number. First, the rule-of-three detector required the third item in a list to be immediately followed by sentence-ending punctuation, so a real list like "labs, classrooms, and offices running smoothly" went undetected because the sentence kept going after it. Second, and more importantly, my entire vocabulary and phrase list had been built from academic and fiction-writing research. It had zero coverage for the specific register AI models default to in business correspondence: "I hope you're all having a wonderful week," "as we navigate," "circle back," "gentle reminder," "sit tight."
After fixing the regex bug and adding a dedicated corporate/email category, the same email now returns nine matched phrases plus the rule-of-three hit it should have caught the first time. That's the difference between a tool that looks sophisticated and one that actually works on the kind of text people write every day.
Try it
Paste in any text and it will flag instances across all seven categories, with the matched phrase highlighted in context. There's a "Load sample" button preloaded with the email described above if you want to see the fix in action. The Help & About panel in the tool documents every category, the sources behind them, and a version changelog, so you can see exactly what changed and why.
One thing the tool is careful to say up front, and worth repeating here: none of this proves authorship. A high score means a passage is worth a second read. A low score doesn't clear it. Every source in this research agrees on that point even when they disagree on everything else.
LLM Advice for System Prompts or Instructions:
## AI WRITING & IMAGE STYLE FILTER — append to system instructions
TEXT OUTPUT RULES
1. Lexical ban list — do not use: delve, underscore (as verb), showcase, leverage, harness, foster, elevate, intricate, meticulous, commendable, seamless, robust, tapestry, landscape (metaphorical), realm, boundaries (metaphorical), "testament to," "plays a vital/pivotal role," "it's worth noting," "in today's fast-paced world."
2. No contrast-framing sentences ("It was not X, it was Y" / "She did not just X, she Y").
3. No three-item lists used as a rhythmic device (e.g., "furious, frightened, undone"). Vary list length; prefer one precise detail over three vague ones.
4. Do not emphasize (bold/italic/intensifier) statements that do not require emphasis. Reserve emphasis for genuinely load-bearing claims.
5. Do not hedge every assertion with a counterpoint. State a position when the evidence supports one.
6. Do not resolve endings or introductions symmetrically or too neatly. Leave appropriate loose ends.
7. Em dash: cap usage near the writer's own historical baseline. Do not use it as a default connector.
8. In narrative/creative text: ground scenes in concrete physical action and sensory detail alongside internal feeling. Do not rely on dialogue and abstract emotion alone.
9. Vary emotional pitch across a scene. Do not sustain one intensity level throughout.
10. Metaphors must survive literal-sense scrutiny. If a comparison cannot be explained plainly, cut it.
IMAGE OUTPUT RULES
1. Hands: render anatomically correct fingers in complex poses (fists, interlaced fingers); verify count and joint plausibility before finalizing.
2. Ears: check for asymmetry or malformation.
3. Background text/signage: must be legible and coherent; no invented or garbled characters.
4. Shadows and reflections: must be consistent with a single light source and physically accurate to the reflected object.
5. Include natural photographic grain/noise; avoid unnaturally smooth or "plastic" textures.
6. Avoid repeating or impossible background architecture/patterns.
REVIEW STEP
Before returning output, self-check against the above list and revise any violation silently.
Sources and further reading
- Human-LLM Coevolution: Evidence from Academic Writing — arXiv, word frequency drop-off after public flagging
- A Longitudinal Study of Word Overuse in Scientific Publications (2020-2024) — the 61-word / 10-significant-word PubMed study
- Delving into PubMed Records: Some Terms in Medical Writing Have Skyrocketed — medRxiv, "delve," "underscore," "meticulous," "commendable" trend data
- Delving into the Utilisation of ChatGPT in Scientific Publications in Astronomy — arXiv, frequency-per-token comparison of AI vs. human text
- Why Does ChatGPT "Delve" So Much? — the 21-focal-word isolation method
- AI words list: the phrases ChatGPT overuses — the 28x/14x/11x frequency multiples
- Let's talk about em dashes in AI — the GPT-3.5 vs. GPT-4o vs. GPT-4.1 controlled generation test and tokenizer explanation
- OpenAI has finally fixed ChatGPT's annoying em dash overuse and OpenAI says it's fixed ChatGPT's em dash problem — the custom-instructions fix, November 2025
- The days of the em dash being a ChatGPT giveaway are over — the "emphasizes everything" observation on why AI text over-flags importance
- How to Spot an AI-Generated Image in 2026: The Tells That Still Work — which 2023-era image tells are now obsolete
- How to Tell if Art is AI-Generated: 8 Signs to Look For (2026) — hands and fingers as the most durable image tell
- AI Image Detection Methods 2026: Tools & Limits — shadow, reflection, and photographic-grain checks
Comments
Post a Comment