Research · 6 min read ·

How often AI tells show up in ChatGPT, GPT-4o, Claude and o1 text

How often do AI tells show up in model text? In 15.98% of ChatGPT answers against 5.93% of human ones, and among 2024 models only GPT-4o clearly stood apart.

GPGenPolishWritten and checked in-house

In short

AI tells, as GenPolish’s detect() defines them, appeared in 15.98% of 26,885 ChatGPT answers and 5.93% of 58,546 human answers to the same questions in the HC3 dataset. In a 2024 news dataset of 150 human articles and 30 per model, only GPT-4o clearly exceeded the journalists, at 53.33% against 21.33%, mostly through clichés. Claude 3.5 Sonnet and o1-pro overlapped with them.

Lists of AI tells circulate everywhere, from em dashes to “delve”, but few of them are counted against human writing on the same prompt. So we ran GenPolish’s own detection engine, the detect() function behind the editor’s marks, over two public datasets in which every model text has a human counterpart written for the same question or the same news subject.

How did we measure AI tells?

The first dataset is HC3, the Human ChatGPT Comparison Corpus, published by Guo and colleagues in January 2023 under a CC-BY-SA-4.0 licence, with each source subset keeping its own licence where that one is stricter. Its English file holds 58,546 human answers and 26,885 ChatGPT answers to the same questions, produced with the GPT-3.5 ChatGPT of December 2022 and January 2023. Reddit’s Explain Like I’m Five forum supplies 51,336 of the human answers.

The second is the Human Detectors dataset released under the MIT licence by Russell, Karpinska and Iyyer for their ACL 2025 paper. It pairs 150 published news and non-fiction articles, mostly from 2024, with 30 articles each from GPT-4o, Claude 3.5 Sonnet and o1-pro, plus two evasion conditions: GPT-4o articles paraphrased afterwards, and o1-pro articles written under a “humaniser” instruction set. The paper itself tested people, and its headline finding is that a majority vote of five frequent ChatGPT users “misclassifies only 1 of 300 articles”.

Each text went through detect() with its default options, and every change it reported was counted under one of the nine categories of our published taxonomy. A text counts as a hit when it contains at least one occurrence, and each share comes with a 95% Wilson confidence interval. Rates per 1,000 words divide all occurrences by all words, a word being any run of characters without whitespace. Two extra markers sit outside the nine tells: the em dash, code point 2014, and the invisible characters our tools look for.

Nothing was sampled. Every non-empty text was measured, and 18 empty ChatGPT answers were set aside. HC3 stores most human answers pre-tokenised, with spaces before punctuation and split contractions, so a minimal detokeniser was applied to every HC3 text, human and machine alike. Human Detectors was measured as published. The aggregate data file is available on request.

How often do AI tells appear in ChatGPT answers?

ChatGPT answers in HC3 contained at least one of the nine tells in 15.98% of cases, with an interval of 15.55 to 16.42. Human answers to the same questions did so in 5.93% of cases, between 5.74 and 6.12. That is about 2.7 times as often, or 1.051 occurrences per 1,000 words against 0.555.

Three categories carry most of that gap. Filler phrases appeared in 4.82% of ChatGPT answers and 0.65% of human ones, padding in 5.73% against 2.09%, and repeated sentence shapes in 5.98% against 2.96%. One category ran the other way: over-signposting turned up in 0.40% of human answers and only 0.05% of ChatGPT answers. Hedges, stock openers and the listicle reflex were close to zero on both sides.

Do GPT-4o, Claude and o1-pro show the same tells?

Human Detectors compares newer models with professional journalists, whose articles contained at least one tell in 21.33% of cases. GPT-4o is the only model whose interval lies clear of that figure: 53.33% of its 30 articles carried a tell. Clichés explain most of it, found in 40.00% of GPT-4o articles and in 2.00% of the human ones.

Share of texts with at least one hit, in per cent, with the 95% Wilson interval in brackets
GroupTextsAny of the nine tellsClichéEm dash
HC3, human answers58,5465.93 (5.74–6.12)0.03 (0.02–0.05)0.59 (0.53–0.65)
HC3, ChatGPT (GPT-3.5)26,88515.98 (15.55–16.42)0.02 (0.01–0.04)0.23 (0.18–0.29)
Human Detectors, journalists15021.33 (15.54–28.56)2.00 (0.68–5.71)82.67 (75.81–87.89)
GPT-4o3053.33 (36.14–69.77)40.00 (24.59–57.68)13.33 (5.31–29.68)
Claude 3.5 Sonnet3036.67 (21.87–54.49)3.33 (0.59–16.67)16.67 (7.34–33.56)
o1-pro3020.00 (9.51–37.31)6.67 (1.85–21.32)100.00 (88.65–100.00)
GPT-4o, paraphrased3033.33 (19.23–51.22)16.67 (7.34–33.56)6.67 (1.85–21.32)
o1-pro, humanised303.33 (0.59–16.67)0.00 (0.00–11.35)100.00 (88.65–100.00)

Claude 3.5 Sonnet reached 36.67% and o1-pro 20.00%, and both intervals overlap the human one. So do the two evasion conditions. Paraphrasing brought GPT-4o down to 33.33%, and the humanised o1-pro articles fell to 3.33%, the lowest of any group. With 30 articles per model, a gap of ten or fifteen points is within the noise, and only the GPT-4o result survives that test.

Is the em dash a sign of AI writing?

The em dash does not mark AI text in either dataset. In HC3, human answers used it more than ChatGPT: 0.59% of human answers contained one against 0.23% of ChatGPT answers, or 0.106 against 0.015 per 1,000 words.

Journalists in Human Detectors are heavy users, with an em dash in 82.67% of their 150 articles. GPT-4o used one in 13.33% of its articles and Claude 3.5 Sonnet in 16.67%, well below them. The exception is o1-pro, which put an em dash in all 30 articles at 8.971 per 1,000 words, about twice the human rate of 4.453. On this evidence the em dash separates one model from journalists, not machines from people.

Do AI texts carry invisible characters?

Invisible characters were rare everywhere. In HC3, 0.19% of human answers and 0.19% of ChatGPT answers contained one. In Human Detectors, one human article out of 150 did, and none of the 150 model articles. Neither dataset shows anything resembling a hidden watermark made of invisible characters.

What can this study not tell you?

  • Model age. HC3 captures ChatGPT as it was in December 2022 and January 2023, and Human Detectors covers GPT-4o, Claude 3.5 Sonnet and o1-pro from 2024 to early 2025. Neither says anything about the models of 2026.
  • Sample size. Human Detectors has 30 texts per model, so its intervals often span about thirty points. Only gaps whose intervals do not overlap are reported as differences here.
  • Narrow domains. HC3 is dominated by Explain Like I’m Five answers written off the cuff, and Human Detectors sets models against edited professional journalism. Neither covers business writing or email.
  • A short, cautious lexicon. detect() holds twenty phrase rules plus two shape checks, so a low rate means those exact phrasings are rare, not that a text has no habits. Its patterns expect a straight apostrophe, and the curly one common in news copy is not matched.
  • Not a detector. A tell found in 16% of ChatGPT answers is also found in 6% of human answers, so none of these markers can say that a given text was machine-written.
  • Pre-tokenised text. The HC3 human answers had to be detokenised. The fix restores punctuation and contractions, but spaced quotes and split hyphens still add a few words to the human denominator.
  • Word definition. A rate per 1,000 words is comparable only with a rate computed the same way, here with words as runs of non-whitespace characters.

What should a writer take from these numbers?

The measurable habits of ChatGPT in 2023 were small verbal ones: filler, padding, repeated sentence shapes. GPT-4o added clichés. Those are worth editing wherever they come from, because each one costs the reader time without adding a fact.

Deleting every em dash or hunting for hidden characters is a different matter. In these datasets neither marker told models from people, and one model used more dashes than the journalists did. A tell is a phrase to rewrite, and a count of tells is a measure of register, never proof of who wrote a page.

Questions

Which AI tell separates ChatGPT answers from human answers the most?

Filler phrases separate ChatGPT answers from human answers the most in the HC3 dataset. GenPolish’s detect() found at least one in 4.82% of 26,885 ChatGPT answers and 0.65% of 58,546 human answers to the same questions. Padding followed at 5.73% against 2.09%, then repeated sentence shapes at 5.98% against 2.96%. Over-signposting went the other way, with 0.40% of human answers against 0.05% of ChatGPT answers.

Do em dashes show that ChatGPT, Claude or o1 wrote a text?

Em dashes do not show that a model wrote a text. In the HC3 dataset, human answers used the em dash more than ChatGPT answers, 0.106 against 0.015 per 1,000 words. In 2024 news articles, 82.67% of 150 human pieces contained one, against 13.33% of 30 GPT-4o and 16.67% of 30 Claude 3.5 Sonnet articles. Only o1-pro used more, in all 30 of its articles.

Can counting AI tells identify a single AI-written text?

Counting AI tells cannot identify a single AI-written text. In the HC3 dataset, 15.98% of ChatGPT answers contained at least one of the nine tells, but so did 5.93% of human answers, and most ChatGPT answers contained none. Trained readers do better: in the Human Detectors study by Russell, Karpinska and Iyyer, five frequent ChatGPT users voting together misclassified only 1 of 300 articles.

Do paraphrased or humanised AI articles show fewer tells?

Humanised o1-pro articles in the Human Detectors dataset showed the fewest tells of any group: 3.33% of 30 articles contained one of the nine, against 21.33% of 150 human articles and 20.00% of 30 plain o1-pro articles. Paraphrased GPT-4o articles fell to 33.33%, from 53.33% for plain GPT-4o. With 30 articles per group, both confidence intervals still overlap the human one.

Sources

  1. HC3: Human ChatGPT Comparison Corpus (dataset card)Hugging Face, Hello-SimpleAI ·
  2. How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and DetectionarXiv ·
  3. human_detectors: data for the ACL 2025 paper on human detection of AI textGitHub, Jenna Russell ·
  4. People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated textarXiv (ACL 2025) ·
Paste a draft and see each tell marked by the same engine used in this study, with every change shown before you keep it.Polish a draft free