TL;DR — How Big Are the Biggest AI Models, and How Sure Are We?
22 frontier models ranked by parameters, from Epoch AI data
Among 22 frontier AI models trained from scratch, Grok 4 and Grok 3 tie for the largest at 3 trillion parameters, and GPT-3 closes the list at 174.6 billion. Epoch AI, whose database supplies the figures, grades each size Confident, Likely or Speculative, and 11 of the 22 fall below Confident, including all of the top five.
The list also leaves out the models that matter most: five of the ten largest training runs, among them GPT-6 Astra, GPT-4.5 and Claude 3.5 Sonnet, have no parameter count on record. It shows the biggest models whose size is known, and one of them, SB-LM from 2007, is an n-gram model rather than a neural network.
! The Three Findings Behind the Ranking:
•
Half the Sizes Are Estimates:11 of 22 figures are graded Likely or Speculative, all of the top five among them. Switch, in sixth, is the first Confirmed bar.
•
The Largest Models Are Missing:Five of the ten biggest training runs have no parameter count on record.
•
A 2007 Model Outsizes GPT-3:SB-LM's 300 billion are n-gram entries, not neural-network weights.
? The Numbers Behind the Story:
•
Largest size listed:3 trillion — Grok 4 and Grok 3, tied
•
Smallest size listed:174.6 billion — GPT-3, 2020
•
Estimated vs confirmed:11 and 11
Parameter count was once the headline number for an AI model. Of the 30 records in Epoch's frontier set published since 2023, 18 carry none, against 1 of 20 in 2020–2022.
Continue reading below for the full detailed article →
Overview
How to Read a Ranking of AI Model Sizes
Parameters are the adjustable weights inside a neural network, and their count is the standard shorthand for a model's size. It measures size, not training effort, which is a separate figure called training compute, recorded in floating-point operations (FLOP). This ranking covers models trained from scratch only: finetunes such as Minerva and Flan-PaLM inherit their parent's count and would repeat PaLM's 540 billion four times. The 22 entries run from 2007 to 2025 and come from Epoch AI's frontier-model database, which admits a model only if it ranked among the five largest training runs of its day.
Three Numbers That Frame the Parameter Ranking
One figure sets the scale of the top bar, one shows how soft the list is, and one shows what is missing.
3 Trillion: The Top Bar Is Rated Speculative
3 trillion
Grok 4 and Grok 3 are both listed at 3 trillion parameters, 17.2 times GPT-3's 174.6 billion. Epoch rates Grok 4's figure Speculative, and the only source in its notes is a rumour of 2.4 trillion. The two bars that set the scale of the whole ranking are estimates rather than measurements.
11 of 22: Sizes That Are Estimates
11 of 22
Epoch grades each size Confident, Likely or Speculative. Eleven of the 22 fall in the bottom two grades, and they include all of the top five. The first Confident figure is Switch at 1.57 trillion, in sixth place. All four open-weight models on the list sit in the confident half, and none of the estimates is open-weight.
5 of 10: Biggest Runs With No Size
5 of 10
Sort Epoch's database by training compute and the ten largest runs include five models with no parameter count on record: GPT-6 Astra, GPT-4.5, Gemini 1.0 Ultra, Grok-2 and Claude 3.5 Sonnet. Whatever the largest AI models weigh, a ranking by parameters cannot see them. It lists the biggest models whose sizes are known.
Grok 4 by the Numbers
The Size: 3 Trillion Parameters 3 trillion (Speculative)Epoch's figure for Grok 4, tied with Grok 3. Its notes cite a rumoured 2.4 trillion instead.
The Effort: About 5e26 FLOP 5e26 FLOP (estimate)Epoch's estimate of Grok 4's training compute, which assumes the same pre-training as Grok 3.
The Gap: 1,600x GPT-3's Compute 1,600 x GPT-3's computeAbout 1,600 times GPT-3's 3.14e23 FLOP, for a model 17.2 times larger in parameters.
What the Ranking Leaves Out
Why the Biggest AI Models Aren't on This List
Sort the same database by training compute instead and the top changes. The ten largest runs are GPT-6 Astra (roughly 1e27 FLOP), Grok 4, GPT-4.5, Grok 3, Llama 4 Behemoth, Gemini 1.0 Ultra, Llama Nemotron Ultra, Llama 3.1-405B, Grok-2 and Claude 3.5 Sonnet. Five of them (GPT-6 Astra, GPT-4.5, Gemini 1.0 Ultra, Grok-2 and Claude 3.5 Sonnet) have no parameter count on record, so no ranking by size can include them.
Some labs chose this. OpenAI's GPT-4 report says it "contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar." Google's PaLM 2 report says details of model size "are withheld from external publication", and its 340 billion figure reached the public through a leak to CNBC.
The database shows the drift: 19 of 20 records published in 2020–2022 carry a parameter count, against 12 of 30 since 2023 (some are successive versions of one model, such as GPT-4o). The pool is narrow too. Epoch's frontier set admits only models in the top five by training compute at release, so DeepSeek-V3 (671 billion total parameters) and Kimi K2 (1 trillion), whose sizes are public, are absent. Either would place in the top eight. The ranking is a record of who still tells us.
Estimated vs Confirmed
How Much of the Top Five Is Actually Known?
Epoch grades each figure on three tiers: Confident (primary or reliable sources giving specific technical details), Likely (the same method with less reliable inputs) and Speculative (a best guess). They correspond to 90% confidence intervals of roughly a factor of 3, 10 and 31. Here Confident is called Confirmed, and Likely and Speculative are called Estimated, a split of exactly 11 and 11. One subtlety: Epoch's grade covers the weakest of three figures, compute, parameters and dataset size, so an Estimated bar is not always a soft parameter count.
The Top Five · 5 of 5 Estimated
Grok 4 and Grok 3 tie at 3 trillion. Epoch rates Grok 4 Speculative, builds its compute estimate on Grok 3's pre-training, and notes a rumour of 2.4 trillion, so the two bars are not independent measurements. Llama 4 Behemoth's roughly 2 trillion rests on Meta's "nearly two trillion total parameters" for a model still in training. GPT-4's 1.8 trillion comes from leaks, a SemiAnalysis report and a separate 1.76 trillion figure, since OpenAI gives no size. Wu Dao 2.0's 1.75 trillion comes from launch coverage and is rated Speculative.
The Confirmed Half · 11 of 22 Sizes
Confirmed bars rest on a paper or on downloadable weights: Switch's paper tabulates 1,571 billion, and GPT-3's appendix lists 174,600 million, not the rounded 175 billion. All four open-weight models (Switch, Llama 3.1-405B, Nemotron-4 340B and Falcon-180B) sit here, and none of the 11 estimates is listed as open-weight. The share shrinks over time: through 2022, 8 of 12 entries are confirmed; from 2023 on, 3 of 10.
The Outlier at #13
Why a 2007 Google Model Out-Sizes GPT-3
SB-LM looks like a data error: a Google model from 2007 with 300 billion parameters, 1.7 times the 174.6 billion of GPT-3, which arrived 13 years later. It is real. The paper is "Large Language Models in Machine Translation" by Thorsten Brants, Ashok Popat, Peng Xu, Franz Och and Jeffrey Dean (EMNLP-CoNLL 2007), whose abstract reports training on up to 2 trillion tokens, "resulting in language models having up to 300 billion n-grams."
An n-gram model is not a neural network. It counts how often short word sequences appear in a corpus and uses those frequencies to score the next word. The paper's largest model, built from 1.8 trillion tokens of web text on 1,500 machines in about a day, holds 300 billion entries in a distributed lookup table, not weights tuned by gradient descent. Its smoothing method, Stupid Backoff (probably the "SB"), carries a name the authors say "originated at a time when we thought that such a simple scheme cannot possibly be good."
Epoch counted the stored n-grams as parameters, so the comparison with GPT-3 is not like for like: SB-LM's estimated 1.4e18 FLOP of training is about 217,000 times less than GPT-3's 3.14e23. It is tagged Estimated even though the 300 billion comes straight from the paper; the soft figure is probably the compute, which assumes a processor. "Parameters" has never meant one thing: weights in a dense network, total experts in a sparse one, entries in a count table.
Dense vs Sparse
Why Total Parameters Flatter Mixture-of-Experts Models
Mixture-of-experts (MoE) models split their parameters into specialist sub-networks and route each input through only a few. Switch's abstract describes "a sparsely-activated model — with outrageous numbers of parameters — but a constant computational cost." A ranking by total parameters credits these models with weights that sit idle for most of any given token.
The gap can be large. Meta calls Llama 4 Behemoth a "288 billion active parameter model with 16 experts" with nearly two trillion in total, so about one parameter in seven works on any token. Epoch's notes describe GPT-4 as a rumoured 1.8 trillion-parameter MoE with 280 billion active, and Wu Dao 2.0 as an MoE with "tens of thousands of experts". With Switch, that is four of the top six bars, ranks 3 to 6. Epoch records no architecture for Grok 3 or Grok 4.
Compute shows it too. Switch's 1.571 trillion parameters took an estimated 8.22e22 FLOP, about a quarter of GPT-3's 3.14e23, with nine times the parameters. PaLM is "densely activated" and Amazon Titan a "200B dense model", so ranking dense PaLM's 540 billion against sparse Switch's 1.571 trillion is not like for like.
When Size Peaked
Why Nine of the 22 Date From 2021
Nine of the 22 models date from 2021, from Switch in January to ERNIE 3.0 Titan on 23 December, against one each in 2020 and 2022. Five of the nine came from Asia: Wu Dao 2.0, Yuan 1.0 and ERNIE 3.0 Titan from China, HyperCLOVA and EXAONE from South Korea. Every China and South Korea entry on the list is from that one year.
Overall, 13 of the 22 are American, three Chinese and two South Korean, with one each from the United Kingdom (Gopher), the UAE (Falcon-180B), Hong Kong (SenseChat) and Israel (Jurassic-1-Jumbo). Eight of the ten entries from 2023 on are American. One year is a thin base for a trend, but it fits the disclosure pattern: after PaLM in 2022, the only confirmed entries are open-weight releases.
Size Is Not Effort
Parameters Versus Training Compute: Why They Diverge
Searches for the "parameters used to train" a model blur two things. A parameter count is a model's size, not a quantity consumed in training; the input that tracks training effort is compute, in FLOP, and the two have diverged. Grok 4 is 17.2 times GPT-3 in parameters but an estimated 1,600 times in compute (5e26 against 3.14e23 FLOP). Llama 3.1-405B has 2.3 times GPT-3's parameters and about 121 times its compute.
GPT-6 Astra, dated 3 September 2026 in the database, is the largest run on record at roughly 1e27 FLOP, twice Grok 4's, within Epoch's range of 5e26 to 2e27. Epoch's notes put it on at least 100,000 GB200 chips, with an estimated power draw of about 233 megawatts. It has no parameter count. Compute is the better-documented axis: Epoch has an estimate for 127 of its 138 frontier models and a parameter count for 107.
All 22 Models Ranked by Parameters, With Epoch's Confidence Grade
Parameter counts are Epoch AI's figures for models trained from scratch. The last two columns give Epoch's grade and the status used here, and the training compute column shows where a model's size and its training effort part ways.
Estimated means Epoch grades the record Likely or Speculative; Confirmed means Confident. Epoch's grade covers the weakest of training compute, parameters and dataset size. Mixture-of-experts entries count total parameters. Ranks 1 and 2 tie at 3 trillion.
Verdict
The Real Ranking Is Who Still Publishes Their Size
The ranking is dependable for one stretch of AI history. In 2020–2022, GPT-3, Gopher, Megatron-Turing NLG, PaLM and Switch reported their sizes in papers, and 19 of the 20 frontier records carried a parameter count. From GPT-3's 174.6 billion to PaLM's 540 billion, those entries are confirmed and comparable, with Switch's sparse 1.571 trillion the one caveat.
The top is where it gives way: the five largest bars are estimates, two tie at 3 trillion, four of the top six are, or are rumoured to be, sparse, and the biggest runs of 2025–2026 are absent. The question for any new largest-model headline is whether the size is published, leaked or inferred.
The data behind this story comes from Epoch AI, via its Data on AI Models database filtered to frontier models, those ranked among the five largest training runs at release. The dataset is published under a Creative Commons Attribution licence, and full credit for collecting, estimating and maintaining it goes to Epoch AI's researchers. Details on SB-LM, Llama 4 Behemoth and GPT-4 were also checked against the original paper, Meta's announcement and OpenAI's report.
FactsFigs reviews, cleans, and cross-checks every source dataset before shaping it into a data story. Each visualization is created and designed in FactsFigs Design Studio — an internal tool developed and owned by FactsFigs — and is the original work of a FactsFigs author, not an AI-generated copy of any existing graphic. Individual assets within a visual may or may not be produced with AI tools, but the design of the visual itself is solely FactsFigs' own.
Figures are estimates at the time of publication, provided for information only — nothing here is financial advice or a guarantee of accuracy.
Last verified: 4 Oct 2026
The charts in this article were built with our own publishing system. See what it does →