Naturalmelo Research · Original Data Study

The AI Vocabulary Gap: What 33,160 Matched Human and AI Texts Reveal About How AI Really Writes

We measured how often 14 famous "AI-sounding" phrases actually appear in matched human and AI writing — 8,290 documents per author, ~16.5 million words. Some AI clichés are up to 205× more frequent in GPT-4o text. Others, including the most famous "AI tell" of all, turned out to be myths.

205×
more GPT-4o documents contain "tapestry" than human documents
26.1%
of GPT-4o texts contain "tapestry" or "delve" — vs 0.36% of human texts
33,160
matched ~500-word documents analyzed (≈16.5M words)
2 / 14
famous "AI tells" that failed the test — including "important to note"

01 — OverviewKey findings

19.8% vs 0.10%"Tapestry" is the strongest single AI marker we measured. It appears in one in five GPT-4o documents but one in a thousand matched human documents — a 205× gap (p < 0.001).
1 in 4 vs 1 in 276The composite check is even sharper. 26.1% of GPT-4o documents contain "tapestry" or "delve"; for human writers it's 0.36% — a 72× gap.
3 fingerprintsModels don't share one style. GPT-4o over-uses tapestry / testament / underscores; GPT-4o mini leads on not only… but also (23.4%); Llama 3 70B prefers in conclusion (11.5% — more than GPT-4o).
0.04% vs 0.21%The most famous "AI tell" failed. "It's important to note" appeared less often in GPT-4o than in human text. "Worth noting" showed no significant difference at all (p = 0.38).

Why it matters: classic connectors like "furthermore" and "moreover" are only 3–7× elevated in AI text yet appear in ~2.5% of genuinely human documents — one reason single-phrase AI spotting produces false accusations. Reliable detection needs base rates, and that's what this study provides.

02 — BackgroundWhy we ran this study

Everyone "knows" that AI writing is full of delve, tapestry, and it's important to note. Teachers circle these phrases; Reddit threads treat them as proof of ChatGPT use; detection tools flag them. But almost nobody has published the base rates: how often do these phrases occur in AI text — and, just as important, how often do real humans use them too?

That second number decides whether a phrase is evidence or a false-accusation machine. So we measured both, using the HAP-E parallel corpus (Reinhart et al., Carnegie Mellon, 2024) — a rare dataset in which humans and LLMs continue the exact same texts, making the comparison apples-to-apples across academic prose, news, fiction, blogs, spoken transcripts, and TV/movie scripts.

03 — The Data14 phrases across 33,160 matched documents

Each cell shows the share of ~500-word documents (n = 8,290 per author) containing the phrase at least once — cell shading is proportional to the rate. Matching was case-insensitive and included word variants (e.g., delve/delving/delves).

Table 1 — Share of documents containing each phrase, by authorHAP-E corpus · n = 8,290 documents ≈ 4.1M words per author · sorted by GPT-4o ÷ Human ratio
PhraseHumanGPT-4oGPT-4o miniLlama 3 70BGPT-4o ÷ Human
tapestry0.10%19.8%20.9%1.5%205×
testament to0.10%16.1%8.4%5.4%167×
in conclusion0.07%7.2%5.4%11.5%100×
underscores / underscoring0.13%12.7%9.9%0.7%96×
vibrant0.17%12.1%16.8%2.3%72×
delve / delving0.27%8.1%8.7%2.9%31×
serves as a0.13%3.3%4.3%0.8%25×
crucial role0.08%2.1%1.7%2.2%24×
not only … but also1.30%15.4%23.4%1.5%12×
showcase / showcasing0.31%3.3%6.2%1.1%11×
moreover2.42%12.3%17.7%6.3%5.1×
furthermore2.48%8.6%11.4%12.6%3.5×
worth noting0.10%0.16%0.24%0.35%1.6× n.s.
important to note0.21%0.04%0.10%0.13%0.2×
Higher share of documents Lower share ≥30× strong AI marker n.s. / reverse not reliable AI evidence

All gaps except "worth noting" are statistically significant at p < 0.01 (chi-square, 2×2, continuity-corrected). Human baseline = the genuine continuation of the same source text (HAP-E "chunk 2").

Figure 1 — The six strongest markers: GPT-4o vs human, same writing tasks
Share of documents containing the phrase · bars scaled to 21%
tapestry
GPT-4o19.8%
Human0.10%
testament to
GPT-4o16.1%
Human0.10%
underscores / underscoring
GPT-4o12.7%
Human0.13%
vibrant
GPT-4o12.1%
Human0.17%
in conclusion
GPT-4o7.2%
Human0.07%
delve / delving
GPT-4o8.1%
Human0.27%
GPT-4o (gpt-4o-2024-08-06)Human (matched continuation)
26.1%
GPT-4o
documents containing "tapestry" or "delve" — more than 1 in 4
72×
the gap between GPT-4o and human writers on the same two-word check — same topics, same length, same genres
0.36%
Human writers
same check, same writing tasks — about 1 in 276 documents

04 — Model ComparisonEvery model has its own fingerprint

A single "AI phrase list" treats all models as one writer. The data says otherwise — each model has a signature set of habits:

GPT-4o

The "tapestry" model — imagery-flavored abstractions
tapestry19.8%
testament to16.1%
underscores12.7%

GPT-4o mini

Smaller model, stronger clichés — structural templates
not only … but also23.4%
tapestry20.9%
moreover17.7%

Llama 3 70B

Avoids "ChatGPT words" — loves essay scaffolding
furthermore12.6%
in conclusion11.5%
tapestry1.5%
Implication: a checker tuned only to 2023-era ChatGPT vocabulary will miss Llama-style output, and vice versa. Reliable AI writing detection has to score many overlapping signals — phrasing, rhythm, structure — not one word list. That is exactly how Naturalmelo's AI checker approaches sentence-level detection.

05 — Myth CheckTwo famous "AI tells" failed the test

Myth busted
"It's important to note…"

Perhaps the most-cited ChatGPT phrase on the internet — yet GPT-4o used it in just 3 of 8,290 documents (0.04%), five times less often than the human writers (0.21%).

Not significant
"Worth noting…"

No statistically significant difference between any model and human writers (p = 0.38). Circling it as AI evidence is guesswork.

Why? Phrase habits are version-specific. Early ChatGPT (GPT-3.5, 2023) hedged constantly in Q&A settings, and the meme stuck. But 2024-generation models writing continuous prose have different habits — the clichés moved from hedging ("important to note") to grandiosity ("a testament to", "a rich tapestry").

Caution for teachers and editors: "furthermore" and "moreover" appear in roughly 2.5% of genuinely human documents — among the highest human base rates of any phrase we tested. Circling a "furthermore" as AI evidence will wrongly accuse real writers far more often than the 205× markers would. Single phrases are probabilistic hints, never proof — a lesson consistent with the ~20% false-positive rates reported industry-wide for AI detectors.

06 — Case Study"Tapestry" invades fiction

The gap is biggest where the word least belongs. When continuing classic fiction — public-domain novels and short stories:

29.1%

of GPT-4o fiction continuations worked in the word "tapestry" (406 of 1,395) — versus 0.5% of the real human continuations (7 of 1,395). A 58× gap inside a genre where the phrase adds nothing.

GPT-4o29.1%
Human0.5%
"The bustling scenes outside the window blurred past, a tapestry of colors and shapes, as adventures waited just beyond the horizon." — GPT-4o, asked to continue a public-domain children's novel (HAP-E, doc fic_0012@gpt-4o-2024-08-06)

This is the signature of template phrasing: the model reaches for the same decorative abstraction regardless of context. It is also why sentence-level review beats a single document score — the flag should point at that sentence, not vaguely at the whole page.

07 — TakeawaysWhat this means for students and writers

If you use AI to draft or polish: the highest-risk phrases are not the ones you've heard about. Before submitting, search your draft for tapestry, testament, underscores, vibrant, delve, showcase, not only… but also — each is 10–205× more common in AI text. Replacing them with concrete, specific wording removes the strongest template signals.

If you write without AI: you are statistically unlikely to trip the strong markers (fewer than 1 in 250 human documents contain "tapestry" or "delve"). But you may well use "furthermore" or "moreover" — and a crude phrase-list detector can flag you anyway. If that happens, version history and drafts remain your best evidence; see our guide on how AI detectors work.

Check your own draft in 30 seconds

Naturalmelo's free AI checker flags template phrasing sentence-by-sentence — including the markers from this study — and rewrites flagged sentences in one click. No login required.

Run a free AI writing check → Or fix flagged text directly with the Clever AI Humanizer

08 — MethodsMethodology

Corpus

We used the Human-AI Parallel Corpus (HAP-E), published by researchers at Carnegie Mellon University (Reinhart, Brown, et al., 2024; arXiv:2410.16107). HAP-E takes 8,290 authentic ~1,000-word human texts across six genres (academic articles, news, fiction, blogs, podcast transcripts, TV/movie scripts), splits each into two ~500-word chunks, gives the first chunk to an LLM, and asks it to write the next ~500 words in the same style and tone. The genuine second human chunk serves as the matched human baseline — same topic, same register, same target length.

Authors compared

Human continuations ("chunk 2"), GPT-4o (gpt-4o-2024-08-06), GPT-4o mini (2024-07-18), and Llama 3 70B Instruct — 8,290 documents each, 33,160 documents and ≈16.5 million words in total. (HAP-E also includes Llama base models, which we excluded as they are not deployed as writing assistants.)

Measurement

For each of 14 phrase patterns we counted the number of documents per author containing the pattern at least once (case-insensitive substring matching; stem patterns capture inflections, e.g. delv- → delve/delves/delving). Counts were obtained with documented SQL (DuckDB) count queries against the corpus via the Hugging Face datasets server on August 2, 2026, and are exactly reproducible. Document-containment share, not raw token frequency, is reported to keep documents of equal weight. Significance: 2×2 chi-square with continuity correction.

Limitations

  • Substring matching can include rare non-target uses (e.g., "underscore" as a character name would count); with n = 8,290 per group such noise is negligible relative to the observed gaps.
  • The corpus normalizes punctuation, so popular punctuation "tells" (like the em dash) could not be tested here.
  • Models are 2024-generation; newer models will shift their fingerprints — that is one of the study's points, and we plan periodic re-runs.
  • English only. Phrase inventories differ by language.
  • Phrase frequency shows how AI tends to write, not whether any single text was AI-written. None of these numbers alone can prove authorship.

Data & license

Findings and figures on this page are released under CC BY 4.0 — reuse freely with attribution to Naturalmelo Research and a link to this page. Underlying corpus © the HAP-E authors (CC BY 4.0).

09 — QuestionsFAQ

Does using "delve" or "tapestry" prove a text was AI-written?

No. These phrases are strong probabilistic markers — "tapestry" is about 205× more common in GPT-4o text than in matched human text — but real humans do use them (about 1 in 1,000 human documents in our data). Treat them as signals to review a sentence, never as proof of authorship.

Is "it's important to note" a reliable sign of ChatGPT?

Not for current models. In our 2026 analysis of 8,290 GPT-4o documents, "important to note" appeared in only 0.04% of them — less often than in matched human writing (0.21%). The phrase was a habit of 2023-era ChatGPT in Q&A settings; modern models writing continuous prose rarely use it.

Which words are the strongest AI markers in 2026?

In this study: tapestry (205×), testament to (167×), in conclusion (100×), underscores (96×), vibrant (72×), and delve (31×), relative to matched human writing. The composite check "contains tapestry OR delve" separates GPT-4o from humans 26.1% vs 0.36%.

Do all AI models overuse the same phrases?

No. GPT-4o favors tapestry/testament/underscores; GPT-4o mini leads on not only… but also and vibrant; Llama 3 70B mostly avoids those but overuses in conclusion and furthermore. Detection based on a single model's word list will miss the others.

Can I be falsely flagged for using "furthermore" or "moreover"?

It's possible with crude phrase-based checking — these connectors appear in ~2.5% of genuinely human documents and are only 3–7× elevated in AI text, the weakest signals we measured. Good detection weighs many signals per sentence instead of keying on one word, and any flag should be reviewable, not an automatic verdict.

How can I check my own draft against these markers?

Naturalmelo's free AI checker scans your draft sentence-by-sentence for template phrasing (including the patterns in this study), shows why each sentence was flagged, and can rewrite flagged sentences with the built-in AI humanizer. It's free, with no login, for drafts of 50–1,000 words.

10 — CitationHow to cite this study

Journalists, teachers, and researchers are welcome to reuse these findings with attribution. Naturalmelo Research (2026). "The AI Vocabulary Gap: Phrase Frequency in 33,160 Matched Human and AI Texts." Naturalmelo. https://www.naturalmelo.com/research/ai-vocabulary-gap (data: HAP-E corpus, Reinhart et al. 2024, arXiv:2410.16107; analysis CC BY 4.0)

Related reading: How AI detectors work · Humanization techniques · Naturalmelo vs Turnitin