Does AI Hallucination Really Get Better With Newer Models?

By Saif Ur Rahman Published 31 Aug 2026 Updated 02 Sept 2026
FactsFigs Data Story
Line Chart
FactsFigs

Facts & Figures

Visual Intelligence

Data Source

TL;DR — Does AI Hallucination Really Get Better With Newer Models?

No. Every provider tested ran flat-to-worse on its newest flagship, and reasoning models make it sharply worse.

On Vectara's grounded-summarization benchmark, OpenAI, Anthropic, Google, xAI and DeepSeek were each tested against the flagship model they replaced — and not one improved meaningfully. Google's Gemini 3 Pro preview scored 13.6% against Gemini 2.5 Pro's 7.0%, and xAI's Grok-4-fast-reasoning scored 20.2% against Grok-3's 5.8% — the two sharpest regressions in the dataset. Reasoning models are the clearest driver: OpenAI's own system card shows o4-mini hallucinating on 48% of PersonQA questions, three times o1's 16%, and the company says it doesn't fully understand why.

! What the Benchmark Data Actually Shows:

  • The Popular Belief: Each new AI model generation hallucinates less than the one before it, as training and safety work improve.
  • What the Flagship Comparison Found: Zero of five providers improved meaningfully on Vectara's benchmark; Google and xAI roughly doubled or tripled their prior flagship's rate.
  • What Reasoning Models Do: OpenAI's and DeepSeek's reasoning-specialized models hallucinate two to three times more than their standard-tier counterparts on the same tests.

? The Numbers Behind the Verdict:

  • Flagship Regression: Google 7.0% to 13.6%; xAI 5.8% to 20.2% (Vectara leaderboard)
  • Reasoning Gap: o3-pro 23.3% vs. GPT-5.4-pro 8.3%; o4-mini-high 18.6% vs. GPT-5.4-mini 5.5% (Vectara)
  • OpenAI's Own Number: PersonQA: o1 16% to o3 33% to o4-mini 48% (OpenAI system card)

The chain of assumptions behind 'newer means safer' doesn't survive contact with benchmark data. Every major provider tested was flat-to-worse on its own newest flagship, and the sharpest driver is reasoning itself: OpenAI's own numbers show its newest reasoning models hallucinating two to three times more than the generation before, with the company stating plainly that it doesn't know why. The one honest complication is that it isn't a straight line — GPT-4.5 improved over GPT-4o right before o3 and o4-mini's regression — so the real answer is 'no, mostly,' not 'no, always.'

Continue reading below for the full detailed article →

Overview

The Myth Doesn't Survive the Benchmark

Every AI lab talks about hallucination the same way: it's a shrinking problem, and each new release is safer than the last. That's become the working assumption every time a flagship model launches, and it comes with an intuitive explanation attached — more training data, more reinforcement learning, more reasoning capacity should mean fewer confident wrong answers. Vectara's Hallucination Leaderboard puts that framing to a direct test: the same grounded-summarization benchmark run across 97 models from every major lab, refreshed against a harder 7,700-document test set in late 2025 so a 2024 flagship and its 2026 successor are scored on identical ground. Run that comparison for OpenAI, Anthropic, Google, xAI and DeepSeek's own newest-versus-previous flagships, and not one improved meaningfully. Two, Google and xAI, got dramatically worse. Narrow the comparison further, to reasoning models specifically, and the gap gets sharper still — OpenAI's own published numbers put newer reasoning models at two to three times the hallucination rate of the generation they replaced, with the company saying outright it doesn't fully understand why. What follows works through the flagship comparison, the reasoning-model gap, the one place the trend isn't a straight line, and the methodology caveat that makes any single hallucination number easy to misquote.

The Three Numbers That Break the 'Newer Is Safer' Myth

Vectara tested 97 models from every major lab on the same grounded-summarization benchmark, and OpenAI separately published its own open-domain PersonQA numbers in its o3/o4-mini system card. Put side by side, three figures carry the whole story: no provider improved on its own newest flagship, reasoning models run up to 3.4x higher than standard ones on the identical test, and OpenAI's own newest reasoning model hallucinates on nearly half of PersonQA's open-domain questions.

Flagship Comparison · 0 of 5 Improved

0/5

OpenAI, Anthropic, Google, xAI and DeepSeek were each tested against the flagship they replaced on Vectara's benchmark. OpenAI and Anthropic held roughly flat. Google's Gemini 3 Pro preview scored 13.6% against Gemini 2.5 Pro's 7.0%, and xAI's Grok-4-fast-reasoning scored 20.2% against Grok-3's 5.8% — nearly doubling and more than tripling their own predecessor's rate.

Reasoning vs. Standard · Up to 3.4x Higher

3.4x

On the identical Vectara benchmark, OpenAI's o3-pro scored 23.3% against GPT-5.4-pro's 8.3%, and o4-mini-high scored 18.6% against GPT-5.4-mini's 5.5% — a 3.4x gap, the widest in the dataset. DeepSeek-R1 scored 11.3% against DeepSeek-V3's 6.1%, nearly double. Not every case moves that much: xAI's Grok-4-fast barely shifts when reasoning is toggled on, 19.7% versus 20.2%.

OpenAI's Own Number · 48% on PersonQA

48%

OpenAI's own o3-and-o4-mini system card reports PersonQA hallucination rates of 16% for o1, 33% for o3, and 48% for o4-mini — roughly double and triple o1's rate in two release cycles. OpenAI attributes it loosely to reasoning models making more claims overall, both correct and hallucinated, and states plainly that the underlying cause is not fully understood.

Reasoning Models, by the Numbers

  • Zero of Five Providers Improved 0/5 OpenAI, Anthropic, Google, xAI and DeepSeek each tested a newest-vs-predecessor flagship pair on Vectara's benchmark. None improved meaningfully.
  • Up to 3.4x Higher for Reasoning Models 3.4x OpenAI's o4-mini-high scored 18.6% against GPT-5.4-mini's 5.5% on the same Vectara test — the widest reasoning-vs-standard gap in the dataset.
  • 48% — o4-mini's PersonQA Rate 48% OpenAI's own newest reasoning model hallucinates on nearly half of PersonQA's open-domain questions, up from o1's 16% two release cycles earlier.

The Assumption

Why 'Newer Means Safer' Became the Assumption

The belief that AI hallucination gets rarer with each new model release didn't come from nowhere. Every major lab frames its launches that way: benchmark charts showing improved factuality scores, release notes citing fewer ungrounded answers, safety sections describing new mitigation techniques. Scale itself seemed to back the story — bigger models, more training data, and reinforcement learning from human feedback all correlated loosely with better performance on most benchmarks through 2023 and 2024, so it was reasonable to extend that line forward and assume hallucination would keep shrinking the same way. Reasoning models arrived on top of that assumption, marketed as the most careful, most deliberate class of model yet — the ones that check their own work before answering. That branding is exactly what makes the actual benchmark numbers, run on the same test across every major provider, worth checking against instead of taking on faith.

The Benchmark

What Vectara's Leaderboard Actually Tests

Vectara's Hallucination Leaderboard is the closest thing to an apples-to-apples test across labs: every model gets the same articles and the same instruction — summarize using only the provided text — and Vectara's own detection model scores each summary for factual consistency, flagging anything below a 0.5 confidence threshold as a hallucination. In November 2025, Vectara overhauled the test set from roughly 1,000 short articles to more than 7,700 longer ones, up to 32,000 tokens, spanning law, medicine, finance, technology and several other domains, and re-scored every model — old and new alike — against the harder set. That re-scoring matters here: the 9.6% carried by GPT-4o and the 13.6% carried by Gemini 3 Pro preview were both measured on the identical current benchmark, not two different eras of test.

The Flagship Comparison

Five Providers, Zero Real Improvements

Run the same before-and-after test on every major lab's newest flagship against the one it replaced, and the results split into two groups rather than the clean downward line the myth predicts. OpenAI's GPT-5.5 (9.3%) and Anthropic's Claude Sonnet 4.6 (10.6%) and Claude Opus 4.7 (12.0%) all landed within half a point of their predecessors — GPT-4o at 9.6%, Sonnet 4 at 10.3%, Opus 4 at 12.0% exactly. Nothing moved enough to call it progress, but nothing collapsed either.

Google and xAI · The Two Real Outliers

Google and xAI are the two providers where the myth doesn't just fail to hold — it inverts. Gemini 3 Pro preview scored 13.6% against Gemini 2.5 Pro's 7.0%, nearly double, and didn't crack Vectara's top 25 despite being Google's newest and most capable model. Grok-4-fast-reasoning scored 20.2% against Grok-3's 5.8%, more than tripling it and landing among the highest hallucination rates of any frontier model tested. Both companies shipped reasoning-heavy successors — which, as the next section shows, is very likely not a coincidence.

DeepSeek sits between the two groups: DeepSeek-V4-Pro scored 8.6% against DeepSeek-V3's 6.1%, a real but smaller regression than Google's or xAI's. Across all five providers checked, the tally is zero meaningfully improved, three held roughly flat or worsened slightly, and two got substantially worse — not the picture 'each generation hallucinates less' predicts.

Reasoning Models

Why Reasoning Models Hallucinate More, Not Less

Reasoning models were supposed to be the safer category — models that think through a problem step by step before answering should, in theory, catch their own errors. The Vectara data says otherwise. OpenAI's o3-pro scored 23.3% against its standard-tier sibling GPT-5.4-pro's 8.3%, and o4-mini-high scored 18.6% against GPT-5.4-mini's 5.5% — a gap of up to 3.4x, the widest reasoning-versus-standard split in the dataset. DeepSeek's R1 scored 11.3% against DeepSeek-V3's 6.1%, nearly double.

OpenAI's Own Numbers Say the Same Thing

OpenAI's own system card for o3 and o4-mini, published alongside the models in April 2025, runs a separate open-domain test called PersonQA and finds the identical pattern on its own benchmark: o1, the previous reasoning generation, hallucinated on 16% of questions, and o3-mini on 14.8%. o3 jumped to 33%, roughly double o1's rate, and o4-mini reached 48%, close to triple it. OpenAI's stated explanation is that reasoning models make more claims overall — more correct ones, but also more hallucinated ones — and the company writes plainly that more research is needed to understand why the rate keeps climbing.

The effect isn't just 'reasoning mode' as a universal on-off switch, either. xAI's Grok-4-fast is the same underlying model tested with reasoning turned off and on, and the two scores are close: 19.7% without reasoning, 20.2% with it. The largest gaps in this dataset all come from comparing separate, reasoning-specialized model lines — OpenAI's o-series, DeepSeek's R1 — against their standard siblings, not from flipping a switch inside one model.

OpenAI's Own System Card

"More Research Is Needed" — OpenAI Still Doesn't Know Why Its Newest Reasoning Models Hallucinate Three Times More Than Its Last.

The Honest Nuance

Why This Isn't a Straight Line Downward

None of this means every new model hallucinates more than the one before it — the trend is genuinely non-monotonic, not a uniform decline. On OpenAI's own PersonQA benchmark, GPT-4.5, a standard non-reasoning model released shortly before o3, scored 19% against GPT-4o's 30%, an improvement of more than a third and the best result in that entire sequence. It's only when o3 and o4-mini launch right after that the numbers jump back up to 33% and 48%. The honest read of the data isn't 'AI hallucination always gets worse' — it's that the recent regression is concentrated specifically in reasoning models, while non-reasoning progress kept moving in the expected direction right up until it didn't.

The Caveat

Why One Model Can Score 0.7% and 48% at Once

Every number in this piece names its benchmark for a reason: there is no single 'hallucination rate' that applies to a model across every kind of task. Vectara's leaderboard measures something narrow and specific — whether a model stays faithful to a document it was told to summarize. PersonQA measures something completely different — whether a model gets open-domain factual questions about real people right without a source document to lean on. Those are different skills, and they produce very different absolute numbers for the same models.

The spread can be dramatic. Gemini-2.0-Flash reportedly scored 0.7% on Vectara's leaderboard but around 7.6% on Vectara's own harder FaithJudge variant — more than ten times higher on the same model, under a harder test. GPT-5 has been reported scoring 47% on the unrelated SimpleQA open-domain benchmark, a number that would look alarming set next to its low single-digit Vectara summarization score, if the two were ever compared directly rather than treated as measuring different things. Quoting a hallucination percentage without naming the benchmark behind it is one of the easiest ways to make AI look better, or worse, than the data actually supports.

Incidents of ai hallucinations

popular incidents of ai hallucinations reported from 2023 to 2025

From fabricated legal citations that got lawyers sanctioned to a $100 billion stock drop over a false telescope claim, these real-world cases show how AI hallucinations moved from harmless quirks to costly, headline-making failures across law, tech, healthcare, and consulting.

Bard's James Webb Telescope error2023~$100 billion in Alphabet market value wiped out (7.7% share drop)Google/Alphabet, investorsBard (Google, powered by LaMDA)
Mata v. Avianca - fabricated legal citations2023$5,000 court sanction (joint, against both lawyers and firm)Roberto Mata (plaintiff), Avianca Airlines (defendant), Steven Schwartz and Peter LoDuca (Levidow, Levidow and Oberman)ChatGPT (OpenAI)
Moffatt v. Air Canada - chatbot bereavement fare2024 (ruling; incident occurred Nov 2022)CA$812.02 (fare difference plus interest plus tribunal fees)Jake Moffatt (customer), Air CanadaAir Canada's website chatbot (unnamed)
AI Overviews 'pizza glue' incident2024No direct financial penalty - reputational damage, PR crisisGoogle, general public/usersGoogle AI Overviews (Gemini-based)
9.11 vs 9.9 math comparison error2024No financial penalty - reputational/credibility damageChatGPT and other LLMs broadly, general publicMultiple models (not one company)
Whisper medical transcription hallucinations2024 (reported)No confirmed financial penalty found - patient-safety risk flagged by researchersOpenAI, Nabla (deployer), 45000+ clinicians / 85+ health systemsWhisper (OpenAI), via Nabla
Deloitte Australia government report2025Partial refund of AUD $440,000 contract (exact amount undisclosed)Australian Dept. of Employment and Workplace Relations, Deloitte Australia, Chris Rudge (Sydney University)Azure OpenAI GPT-4o (via Deloitte)

The Verdict

So, Does AI Hallucination Get Better? Mostly No.

Run the numbers the way the myth predicts — pick any provider's newest flagship, compare it to the one it replaced, expect a lower hallucination rate — and it works for exactly zero of the five providers checked here. OpenAI and Anthropic held flat. DeepSeek got slightly worse. Google nearly doubled its rate and xAI more than tripled it. On the benchmark built specifically to make this comparison fair, 'newer is safer' does not survive contact with the data.

The sharper story sits inside that result, not around it: reasoning is the real driver. Strip out everything except reasoning-versus-standard pairs from the same provider, and the gap runs up to 3.4x higher, confirmed independently by OpenAI's own published system card. That's the part of this myth-bust that should worry anyone deploying these models for anything involving names, dates, or citations — and it's the part OpenAI itself says it can't yet explain.

The one thing this data doesn't support is the flip side of the myth — that AI hallucination is now uniformly worse, full stop. GPT-4.5 improved. Anthropic's flagships barely moved either direction. What changed is specifically tied to reasoning-model releases, on specific benchmarks, and any claim that skips naming which benchmark it's quoting is telling less than the full story.

Data Source and Attribution

Vectara Hallucination Leaderboard OpenAI o3 and o4-mini System Card TechCrunch

The data behind this story comes primarily from Vectara's Hallucination Leaderboard, covering 105 models scored on an identical grounded-summarization benchmark, and from OpenAI's own o3 and o4-mini system card, which publishes PersonQA hallucination rates directly. Additional context and cross-benchmark figures were drawn from TechCrunch's reporting on the o3/o4-mini system card, Yahoo/Mashable's coverage of the GPT-4.5-to-GPT-4o comparison, and Seekr's and TrueStandard's write-ups on benchmark methodology. Full credit for building and maintaining the leaderboard goes to Vectara; full credit for the system card goes to OpenAI.

FactsFigs reviews, cleans, and cross-checks every source dataset before shaping it into a data story. Each visualization is created and designed in FactsFigs Design Studio — an internal tool developed and owned by FactsFigs — and is the original work of a FactsFigs author, not an AI-generated copy of any existing graphic. Individual assets within a visual may or may not be produced with AI tools, but the design of the visual itself is solely FactsFigs' own.

Figures are estimates at the time of publication, provided for information only — nothing here is financial advice or a guarantee of accuracy.

Last verified: 31 Aug 2026

The charts in this article were built with our own publishing system. See what it does →