The AI Vocabulary Gap: What 33,160 Matched Human and AI Texts Reveal About How AI Really Writes
We measured how often 14 famous "AI-sounding" phrases actually appear in matched human and AI writing — 8,290 documents per author, ~16.5 million words. Some AI clichés are up to 205× more frequent in GPT-4o text. Others, including the most famous "AI tell" of all, turned out to be myths.
01 — OverviewKey findings
Why it matters: classic connectors like "furthermore" and "moreover" are only 3–7× elevated in AI text yet appear in ~2.5% of genuinely human documents — one reason single-phrase AI spotting produces false accusations. Reliable detection needs base rates, and that's what this study provides.
02 — BackgroundWhy we ran this study
Everyone "knows" that AI writing is full of delve, tapestry, and it's important to note. Teachers circle these phrases; Reddit threads treat them as proof of ChatGPT use; detection tools flag them. But almost nobody has published the base rates: how often do these phrases occur in AI text — and, just as important, how often do real humans use them too?
That second number decides whether a phrase is evidence or a false-accusation machine. So we measured both, using the HAP-E parallel corpus (Reinhart et al., Carnegie Mellon, 2024) — a rare dataset in which humans and LLMs continue the exact same texts, making the comparison apples-to-apples across academic prose, news, fiction, blogs, spoken transcripts, and TV/movie scripts.
03 — The Data14 phrases across 33,160 matched documents
Each cell shows the share of ~500-word documents (n = 8,290 per author) containing the phrase at least once — cell shading is proportional to the rate. Matching was case-insensitive and included word variants (e.g., delve/delving/delves).
| Phrase | Human | GPT-4o | GPT-4o mini | Llama 3 70B | GPT-4o ÷ Human |
|---|---|---|---|---|---|
| tapestry | 0.10% | 19.8% | 20.9% | 1.5% | 205× |
| testament to | 0.10% | 16.1% | 8.4% | 5.4% | 167× |
| in conclusion | 0.07% | 7.2% | 5.4% | 11.5% | 100× |
| underscores / underscoring | 0.13% | 12.7% | 9.9% | 0.7% | 96× |
| vibrant | 0.17% | 12.1% | 16.8% | 2.3% | 72× |
| delve / delving | 0.27% | 8.1% | 8.7% | 2.9% | 31× |
| serves as a | 0.13% | 3.3% | 4.3% | 0.8% | 25× |
| crucial role | 0.08% | 2.1% | 1.7% | 2.2% | 24× |
| not only … but also | 1.30% | 15.4% | 23.4% | 1.5% | 12× |
| showcase / showcasing | 0.31% | 3.3% | 6.2% | 1.1% | 11× |
| moreover | 2.42% | 12.3% | 17.7% | 6.3% | 5.1× |
| furthermore | 2.48% | 8.6% | 11.4% | 12.6% | 3.5× |
| worth noting | 0.10% | 0.16% | 0.24% | 0.35% | 1.6× n.s. |
| important to note | 0.21% | 0.04% | 0.10% | 0.13% | 0.2× |
All gaps except "worth noting" are statistically significant at p < 0.01 (chi-square, 2×2, continuity-corrected). Human baseline = the genuine continuation of the same source text (HAP-E "chunk 2").
04 — Model ComparisonEvery model has its own fingerprint
A single "AI phrase list" treats all models as one writer. The data says otherwise — each model has a signature set of habits:
GPT-4o
GPT-4o mini
Llama 3 70B
05 — Myth CheckTwo famous "AI tells" failed the test
Perhaps the most-cited ChatGPT phrase on the internet — yet GPT-4o used it in just 3 of 8,290 documents (0.04%), five times less often than the human writers (0.21%).
No statistically significant difference between any model and human writers (p = 0.38). Circling it as AI evidence is guesswork.
Why? Phrase habits are version-specific. Early ChatGPT (GPT-3.5, 2023) hedged constantly in Q&A settings, and the meme stuck. But 2024-generation models writing continuous prose have different habits — the clichés moved from hedging ("important to note") to grandiosity ("a testament to", "a rich tapestry").
06 — Case Study"Tapestry" invades fiction
The gap is biggest where the word least belongs. When continuing classic fiction — public-domain novels and short stories:
of GPT-4o fiction continuations worked in the word "tapestry" (406 of 1,395) — versus 0.5% of the real human continuations (7 of 1,395). A 58× gap inside a genre where the phrase adds nothing.
"The bustling scenes outside the window blurred past, a tapestry of colors and shapes, as adventures waited just beyond the horizon." — GPT-4o, asked to continue a public-domain children's novel (HAP-E, doc fic_0012@gpt-4o-2024-08-06)
This is the signature of template phrasing: the model reaches for the same decorative abstraction regardless of context. It is also why sentence-level review beats a single document score — the flag should point at that sentence, not vaguely at the whole page.
07 — TakeawaysWhat this means for students and writers
If you use AI to draft or polish: the highest-risk phrases are not the ones you've heard about. Before submitting, search your draft for tapestry, testament, underscores, vibrant, delve, showcase, not only… but also — each is 10–205× more common in AI text. Replacing them with concrete, specific wording removes the strongest template signals.
If you write without AI: you are statistically unlikely to trip the strong markers (fewer than 1 in 250 human documents contain "tapestry" or "delve"). But you may well use "furthermore" or "moreover" — and a crude phrase-list detector can flag you anyway. If that happens, version history and drafts remain your best evidence; see our guide on how AI detectors work.
Check your own draft in 30 seconds
Naturalmelo's free AI checker flags template phrasing sentence-by-sentence — including the markers from this study — and rewrites flagged sentences in one click. No login required.
Run a free AI writing check → Or fix flagged text directly with the Clever AI Humanizer08 — MethodsMethodology
Corpus
We used the Human-AI Parallel Corpus (HAP-E), published by researchers at Carnegie Mellon University (Reinhart, Brown, et al., 2024; arXiv:2410.16107). HAP-E takes 8,290 authentic ~1,000-word human texts across six genres (academic articles, news, fiction, blogs, podcast transcripts, TV/movie scripts), splits each into two ~500-word chunks, gives the first chunk to an LLM, and asks it to write the next ~500 words in the same style and tone. The genuine second human chunk serves as the matched human baseline — same topic, same register, same target length.
Authors compared
Human continuations ("chunk 2"), GPT-4o (gpt-4o-2024-08-06), GPT-4o mini (2024-07-18), and Llama 3 70B Instruct — 8,290 documents each, 33,160 documents and ≈16.5 million words in total. (HAP-E also includes Llama base models, which we excluded as they are not deployed as writing assistants.)
Measurement
For each of 14 phrase patterns we counted the number of documents per author containing the pattern at least once (case-insensitive substring matching; stem patterns capture inflections, e.g. delv- → delve/delves/delving). Counts were obtained with documented SQL (DuckDB) count queries against the corpus via the Hugging Face datasets server on August 2, 2026, and are exactly reproducible. Document-containment share, not raw token frequency, is reported to keep documents of equal weight. Significance: 2×2 chi-square with continuity correction.
Limitations
- Substring matching can include rare non-target uses (e.g., "underscore" as a character name would count); with n = 8,290 per group such noise is negligible relative to the observed gaps.
- The corpus normalizes punctuation, so popular punctuation "tells" (like the em dash) could not be tested here.
- Models are 2024-generation; newer models will shift their fingerprints — that is one of the study's points, and we plan periodic re-runs.
- English only. Phrase inventories differ by language.
- Phrase frequency shows how AI tends to write, not whether any single text was AI-written. None of these numbers alone can prove authorship.
Data & license
Findings and figures on this page are released under CC BY 4.0 — reuse freely with attribution to Naturalmelo Research and a link to this page. Underlying corpus © the HAP-E authors (CC BY 4.0).
09 — QuestionsFAQ
Does using "delve" or "tapestry" prove a text was AI-written?
No. These phrases are strong probabilistic markers — "tapestry" is about 205× more common in GPT-4o text than in matched human text — but real humans do use them (about 1 in 1,000 human documents in our data). Treat them as signals to review a sentence, never as proof of authorship.
Is "it's important to note" a reliable sign of ChatGPT?
Not for current models. In our 2026 analysis of 8,290 GPT-4o documents, "important to note" appeared in only 0.04% of them — less often than in matched human writing (0.21%). The phrase was a habit of 2023-era ChatGPT in Q&A settings; modern models writing continuous prose rarely use it.
Which words are the strongest AI markers in 2026?
In this study: tapestry (205×), testament to (167×), in conclusion (100×), underscores (96×), vibrant (72×), and delve (31×), relative to matched human writing. The composite check "contains tapestry OR delve" separates GPT-4o from humans 26.1% vs 0.36%.
Do all AI models overuse the same phrases?
No. GPT-4o favors tapestry/testament/underscores; GPT-4o mini leads on not only… but also and vibrant; Llama 3 70B mostly avoids those but overuses in conclusion and furthermore. Detection based on a single model's word list will miss the others.
Can I be falsely flagged for using "furthermore" or "moreover"?
It's possible with crude phrase-based checking — these connectors appear in ~2.5% of genuinely human documents and are only 3–7× elevated in AI text, the weakest signals we measured. Good detection weighs many signals per sentence instead of keying on one word, and any flag should be reviewable, not an automatic verdict.
How can I check my own draft against these markers?
Naturalmelo's free AI checker scans your draft sentence-by-sentence for template phrasing (including the patterns in this study), shows why each sentence was flagged, and can rewrite flagged sentences with the built-in AI humanizer. It's free, with no login, for drafts of 50–1,000 words.
10 — CitationHow to cite this study
Naturalmelo Research (2026). "The AI Vocabulary Gap: Phrase Frequency in 33,160 Matched Human and AI Texts." Naturalmelo. https://www.naturalmelo.com/research/ai-vocabulary-gap (data: HAP-E corpus, Reinhart et al. 2024, arXiv:2410.16107; analysis CC BY 4.0)
Related reading: How AI detectors work · Humanization techniques · Naturalmelo vs Turnitin