How AI Detectors Spot AI Writing: Four Tells, 2023-2026

By Saif Ur Rahman Published 15 Sept 2026 Updated 15 Sept 2026

TL;DR — How Does an AI Detector Spot AI Writing?

Not by understanding the text. By counting habits — word choice, punctuation, sentence shape — that models use far more often than people do.

Pew Research Center ran an AI detection tool called Open Pangram, built by Pangram, over almost half a million English-language webpages from the Common Crawl archive. In the July 2026 snapshot, 10% of all sampled pages showed significant signs of AI authorship. The same analysis tracked four of the habits these models key on. Since January 2023 the em dash has risen 93% per 10,000 words, Oxford commas 63%, a basket of AI-typical vocabulary 118%, and negative parallelism — the "it's not just X, it's Y" construction — 171%. Three of those four peaked in mid-2025 and have since fallen. The em dash is the only one still accelerating, and it spent the whole of 2023 going the other way.

! What the Detector Is Actually Counting:

  • The Method: A machine learning model trained to spot words, phrases and linguistic quirks used more often by AI than by human authors — pattern frequency, not meaning.
  • The Four Markers Tracked: Em dashes, Oxford commas, a 27-word AI vocabulary list, and negative parallelism. All four are more common on the web now than in January 2023.
  • The Catch: Every marker is a population statistic. A human writer who favours em dashes produces the same signal as a language model.

Continue reading below for the full detailed article →

Overview

A Detector Reads Habits, Not Meaning

An AI detector does not read a page and judge whether a machine could have written it. It counts things. Specific words, specific punctuation, specific sentence shapes — and it compares how often they appear against how often a person would use them. That is why detection works at all, and also why it goes wrong in the particular way it does. The signal is real: language models genuinely over-use certain constructions, and Pew Research Center's analysis of 490,000 webpages shows those constructions spreading across the open web year after year. But a frequency is a property of a population, not of a paragraph. The four markers below are the clearest public evidence of what a detector is looking at — and of why a single flagged document proves much less than the number attached to it suggests.

The Method

What the Detector Is Actually Measuring

Pew used Open Pangram, a detection tool built by the company Pangram, which the researchers describe as a machine learning model that "looks at patterns in language, identifying words, phrases and linguistic quirks that are more commonly used by AI than by human authors." Nothing in that description involves comprehension. The model has learned which surface features correlate with machine authorship, and it reports how strongly a document carries them.

The sample was almost half a million English-language webpages drawn from Common Crawl, a public archive of the open web, covering January 2021 through July 2026. For the marker series specifically, Pew filtered to pages published after 30 November 2022 — the day ChatGPT was released — and plotted each point as a six-month average, which smooths the noise a single crawl would introduce.

The Tells

The Four Habits Being Counted

Two of the markers are punctuation. The em dash is the notorious one, to the point that writers now strip it out of their own drafts to avoid suspicion. The Oxford comma — the comma before "and" in a list — is the quieter one, and at 34 to 56 uses per 10,000 words it is by far the most common of the four.

The third is vocabulary: a 27-word basket Pew treats as AI-typical, including "delve", "tapestry", "testament", "intricate", "meticulous", "pivotal", "showcase" and "underscore". None is unusual on its own. Density is the signal.

The fourth is structural, and the most interesting. Negative parallelism is the "it's not just X, it's Y" construction — a rhetorical move language models reach for constantly when asked to sound emphatic. At under three uses per 10,000 words it is the rarest marker Pew tracks, and the one a human writer is least likely to produce repeatedly by accident.

The Outlier

The Em Dash Spent Two Years Falling First

If the em dash were simply an AI fingerprint, its frequency should have climbed from the moment ChatGPT arrived. It did the opposite. From 5.79 uses per 10,000 words in January 2023 it dropped to 4.2 that July, recovered to 5.32, fell again to 4.58, and only passed its 2023 starting point in January 2025.

Then it doubled in a year, reaching 11.19 by January 2026. On the indexed view it ends at 193 against the Oxford comma's 163 — the steepest final climb of the four despite the worst start.

The lag matters for anyone reading a detector score. The marker most strongly associated with AI writing in public argument was, for the first two years of the ChatGPT era, a marker that was becoming less common. Whatever the em dash now measures, it did not start measuring it in 2022.

The Turn

Three of the Four Have Started Falling

The most recent data point breaks the story's simple version. AI vocabulary peaked at 28.15 uses per 10,000 words in July 2025 and fell to 26.02 by January 2026. Negative parallelism peaked at 2.72 and fell to 2.36. On the indexed view, negative parallelism has come down from 313 to 271.

Only the Oxford comma and the em dash were still rising at the end of the series, and only the em dash was rising sharply. Three markers up overall, three of four down in the final six months, one accelerating — that is a less tidy shape than "AI writing is taking over the web", and it is what the data shows.

Pew does not offer a cause, and neither will this piece. A drop in "delve" could mean models changed, or writers learned to edit it out, or the mix of pages in the sample shifted. The series measures the web's output, not any single model's behaviour.

The Catch

Why a Frequency Cannot Judge One Page

Every number here is a rate across hundreds of thousands of documents. That is what makes the trend trustworthy and what makes it useless as proof about any individual piece of writing. A person who has used em dashes for thirty years produces exactly the same signal as a model that picked the habit up last year.

Pew is explicit about the limitation: "AI detection models aren't perfect — they sometimes misclassify individual documents that were written by humans as including signs of AI authorship, and vice versa." That sentence is doing a lot of work. It is the difference between a measurement and an accusation.

There is a second limitation built into the method. Every marker here is public knowledge, which means every marker is avoidable. A writer who wants to defeat detection needs only to delete their em dashes and rephrase one construction — and the habits that survive that edit are, by definition, the ones the detector was not counting.

The full series

All Four Markers, Per 10,000 Words

Raw frequency at each six-month point, with the indexed value against January 2023 in brackets. The absolute columns show why the chart is indexed: Oxford commas outnumber negative parallelism by more than twentyfold at every point in the series.

Jan 20235.79 (100)34.04 (100)11.94 (100)0.87 (100)
Jul 20234.20 (73)42.57 (125)21.42 (179)1.88 (216)
Jan 20245.32 (92)46.25 (136)25.66 (215)1.90 (218)
Jul 20244.58 (79)44.06 (129)23.56 (197)2.08 (239)
Jan 20256.30 (109)48.35 (142)24.79 (208)2.52 (290)
Jul 20258.58 (148)51.08 (150)28.15 (236)2.72 (313)
Jan 202611.19 (193)55.51 (163)26.02 (218)2.36 (271)

The Verdict

Habits Are Evidence of a Trend, Not of Authorship

What Pew's series establishes is that the fingerprints are real and spreading. Four separate habits, measured the same way across three years and half a million pages, all sit higher now than they did in January 2023. A detector keying on them is not reading tea leaves.

What it cannot establish is that any given page carrying those habits was machine-written. The markers work in aggregate because rare things become common; they fail individually for the same reason, since one page cannot have a frequency. Ten per cent of the web showing AI signals is a finding. This page showing AI signals is a hypothesis.

The useful way to read a detector score, then, is as a question rather than a verdict — and to notice that the strongest tell in the dataset is a punctuation mark that spent its first two post-ChatGPT years becoming rarer.

Data Source and Attribution

Pew Research Center

The data behind this story comes from Pew Research Center's Data Labs analysis, "How much of the internet is written with AI?", published 20 August 2026. Pew sampled almost half a million English-language webpage texts from the Common Crawl web repository and classified them using Open Pangram, an AI detection tool built by Pangram. Marker frequencies are plotted as six-month averages over pages published after 30 November 2022. Full credit for collecting, classifying and publishing the dataset goes to Pew Research Center.

FactsFigs reviews, cleans, and cross-checks every source dataset before shaping it into a data story. Each visualization is created and designed in FactsFigs Design Studio — an internal tool developed and owned by FactsFigs — and is the original work of a FactsFigs author, not an AI-generated copy of any existing graphic. Individual assets within a visual may or may not be produced with AI tools, but the design of the visual itself is solely FactsFigs' own.

Figures are estimates at the time of publication, provided for information only — nothing here is financial advice or a guarantee of accuracy.

The charts in this article were built with our own publishing system. See what it does →