How AI Detectors Spot AI Writing: Four Tells, 2023-2026
Facts & Figures
Visual Intelligence
TL;DR — How Does an AI Detector Spot AI Writing?
Not by understanding the text. By counting habits — word choice, punctuation, sentence shape — that models use far more often than people do.
Pew Research Center ran an AI detection tool called Open Pangram, built by Pangram, over almost half a million English-language webpages from the Common Crawl archive. In the July 2026 snapshot, 10% of all sampled pages showed significant signs of AI authorship. The same analysis tracked four of the habits these models key on. Since January 2023 the em dash has risen 93% per 10,000 words, Oxford commas 63%, a basket of AI-typical vocabulary 118%, and negative parallelism — the "it's not just X, it's Y" construction — 171%. Three of those four peaked in mid-2025 and have since fallen. The em dash is the only one still accelerating, and it spent the whole of 2023 going the other way.
! What the Detector Is Actually Counting:
- • The Method: A machine learning model trained to spot words, phrases and linguistic quirks used more often by AI than by human authors — pattern frequency, not meaning.
- • The Four Markers Tracked: Em dashes, Oxford commas, a 27-word AI vocabulary list, and negative parallelism. All four are more common on the web now than in January 2023.
- • The Catch: Every marker is a population statistic. A human writer who favours em dashes produces the same signal as a language model.
Continue reading below for the full detailed article →
Overview
A Detector Reads Habits, Not Meaning
An AI detector does not read a page and judge whether a machine could have written it. It counts things. Specific words, specific punctuation, specific sentence shapes — and it compares how often they appear against how often a person would use them. That is why detection works at all, and also why it goes wrong in the particular way it does. The signal is real: language models genuinely over-use certain constructions, and Pew Research Center's analysis of 490,000 webpages shows those constructions spreading across the open web year after year. But a frequency is a property of a population, not of a paragraph. The four markers below are the clearest public evidence of what a detector is looking at — and of why a single flagged document proves much less than the number attached to it suggests.
The Method
What the Detector Is Actually Measuring
Pew used Open Pangram, a detection tool built by the company Pangram, which the researchers describe as a machine learning model that "looks at patterns in language, identifying words, phrases and linguistic quirks that are more commonly used by AI than by human authors." Nothing in that description involves comprehension. The model has learned which surface features correlate with machine authorship, and it reports how strongly a document carries them.
The sample was almost half a million English-language webpages drawn from Common Crawl, a public archive of the open web, covering January 2021 through July 2026. For the marker series specifically, Pew filtered to pages published after 30 November 2022 — the day ChatGPT was released — and plotted each point as a six-month average, which smooths the noise a single crawl would introduce.
The Tells
The Four Habits Being Counted
Two of the markers are punctuation. The em dash is the notorious one, to the point that writers now strip it out of their own drafts to avoid suspicion. The Oxford comma — the comma before "and" in a list — is the quieter one, and at 34 to 56 uses per 10,000 words it is by far the most common of the four.
The third is vocabulary: a 27-word basket Pew treats as AI-typical, including "delve", "tapestry", "testament", "intricate", "meticulous", "pivotal", "showcase" and "underscore". None is unusual on its own. Density is the signal.
The fourth is structural, and the most interesting. Negative parallelism is the "it's not just X, it's Y" construction — a rhetorical move language models reach for constantly when asked to sound emphatic. At under three uses per 10,000 words it is the rarest marker Pew tracks, and the one a human writer is least likely to produce repeatedly by accident.
The Outlier
The Em Dash Spent Two Years Falling First
If the em dash were simply an AI fingerprint, its frequency should have climbed from the moment ChatGPT arrived. It did the opposite. From 5.79 uses per 10,000 words in January 2023 it dropped to 4.2 that July, recovered to 5.32, fell again to 4.58, and only passed its 2023 starting point in January 2025.
Then it doubled in a year, reaching 11.19 by January 2026. On the indexed view it ends at 193 against the Oxford comma's 163 — the steepest final climb of the four despite the worst start.
The lag matters for anyone reading a detector score. The marker most strongly associated with AI writing in public argument was, for the first two years of the ChatGPT era, a marker that was becoming less common. Whatever the em dash now measures, it did not start measuring it in 2022.
The Turn
Three of the Four Have Started Falling
The most recent data point breaks the story's simple version. AI vocabulary peaked at 28.15 uses per 10,000 words in July 2025 and fell to 26.02 by January 2026. Negative parallelism peaked at 2.72 and fell to 2.36. On the indexed view, negative parallelism has come down from 313 to 271.
Only the Oxford comma and the em dash were still rising at the end of the series, and only the em dash was rising sharply. Three markers up overall, three of four down in the final six months, one accelerating — that is a less tidy shape than "AI writing is taking over the web", and it is what the data shows.
Pew does not offer a cause, and neither will this piece. A drop in "delve" could mean models changed, or writers learned to edit it out, or the mix of pages in the sample shifted. The series measures the web's output, not any single model's behaviour.
The Catch
Why a Frequency Cannot Judge One Page
Every number here is a rate across hundreds of thousands of documents. That is what makes the trend trustworthy and what makes it useless as proof about any individual piece of writing. A person who has used em dashes for thirty years produces exactly the same signal as a model that picked the habit up last year.
Pew is explicit about the limitation: "AI detection models aren't perfect — they sometimes misclassify individual documents that were written by humans as including signs of AI authorship, and vice versa." That sentence is doing a lot of work. It is the difference between a measurement and an accusation.
There is a second limitation built into the method. Every marker here is public knowledge, which means every marker is avoidable. A writer who wants to defeat detection needs only to delete their em dashes and rephrase one construction — and the habits that survive that edit are, by definition, the ones the detector was not counting.
The full series
All Four Markers, Per 10,000 Words
Raw frequency at each six-month point, with the indexed value against January 2023 in brackets. The absolute columns show why the chart is indexed: Oxford commas outnumber negative parallelism by more than twentyfold at every point in the series.
| Jan 2023 | 5.79 (100) | 34.04 (100) | 11.94 (100) | 0.87 (100) |
| Jul 2023 | 4.20 (73) | 42.57 (125) | 21.42 (179) | 1.88 (216) |
| Jan 2024 | 5.32 (92) | 46.25 (136) | 25.66 (215) | 1.90 (218) |
| Jul 2024 | 4.58 (79) | 44.06 (129) | 23.56 (197) | 2.08 (239) |
| Jan 2025 | 6.30 (109) | 48.35 (142) | 24.79 (208) | 2.52 (290) |
| Jul 2025 | 8.58 (148) | 51.08 (150) | 28.15 (236) | 2.72 (313) |
| Jan 2026 | 11.19 (193) | 55.51 (163) | 26.02 (218) | 2.36 (271) |
The Verdict
Habits Are Evidence of a Trend, Not of Authorship
What Pew's series establishes is that the fingerprints are real and spreading. Four separate habits, measured the same way across three years and half a million pages, all sit higher now than they did in January 2023. A detector keying on them is not reading tea leaves.
What it cannot establish is that any given page carrying those habits was machine-written. The markers work in aggregate because rare things become common; they fail individually for the same reason, since one page cannot have a frequency. Ten per cent of the web showing AI signals is a finding. This page showing AI signals is a hypothesis.
The useful way to read a detector score, then, is as a question rather than a verdict — and to notice that the strongest tell in the dataset is a punctuation mark that spent its first two post-ChatGPT years becoming rarer.
Data Source and Attribution
The data behind this story comes from Pew Research Center's Data Labs analysis, "How much of the internet is written with AI?", published 20 August 2026. Pew sampled almost half a million English-language webpage texts from the Common Crawl web repository and classified them using Open Pangram, an AI detection tool built by Pangram. Marker frequencies are plotted as six-month averages over pages published after 30 November 2022. Full credit for collecting, classifying and publishing the dataset goes to Pew Research Center.
FactsFigs reviews, cleans, and cross-checks every source dataset before shaping it into a data story. Each visualization is created and designed in FactsFigs Design Studio — an internal tool developed and owned by FactsFigs — and is the original work of a FactsFigs author, not an AI-generated copy of any existing graphic. Individual assets within a visual may or may not be produced with AI tools, but the design of the visual itself is solely FactsFigs' own.
Figures are estimates at the time of publication, provided for information only — nothing here is financial advice or a guarantee of accuracy.
Related Publications
More from FactsFigs
Does AI Hallucination Really Get Better With Newer Models?
On Vectara's benchmark, every major AI provider was flat-to-worse on their newest flagship — and reasoning models score up to 3.4x higher.
31 Aug 2026
AI Data Center Electricity Use : Demand Explained
AI data centers could nearly double global electricity demand by 2030 — and bills are already rising where that growth concentrates.
31 Aug 2026
Is DeepSeek Really Winning Countries from ChatGPT ?
A categorical map of 149 countries and territories shows DeepSeek's rare wins are almost always government bans or OpenAI exclusions, not real consumer preference.
20 Aug 2026
How Much Longer Can AI Work Alone in 2026 Than in 2019?
AI's task-completion time horizon grew from a 2-second sentence in 2019 to a 12-hour unsupervised workday by 2026, milestone by milestone.
20 Aug 2026
