LLM API Pricing Compared: Cost per 1M Tokens

By Saif Ur Rahman Published 10 Jul 2026 Updated 17 Jul 2026
FactsFigs Data Story
13 LLM APIs Ranked by Cost per 1M Tokens
FactsFigs

Facts & Figures

Visual Intelligence

TL;DR — What Top LLM APIs Cost per 1M Tokens

LLM API pricing at a glance

This comparison ranks 13 LLM API models from OpenAI, Anthropic, Google, xAI, and DeepSeek by their combined cost per 1M input plus 1M output tokens. The headline finding: pricing spans from $35 at the top (GPT-5.5) to $0.42 at the bottom (DeepSeek V4 Flash) — a roughly 83x spread that makes model selection one of the biggest cost levers in AI product development.

! Key takeaways:

  • GPT-5.5 leads on price: At $35 combined ($5 input / $30 output per 1M tokens), OpenAI's flagship is the most expensive model in the comparison.
  • Claude Opus 4.8 sits close behind: Anthropic's flagship matches GPT-5.5 on input ($5/1M) but undercuts it on output ($25 vs $30), landing at $30 combined.
  • Budget tiers collapse the floor: DeepSeek V4 Flash ($0.42 combined), Gemini 2.5 Flash-lite ($0.50), and GPT-5.4-nano ($1.45) run high-volume workloads at prices 24x to 83x below flagship rates.
  • Output premiums vary by provider: Every provider charges more to generate than to read: 2x input at xAI and DeepSeek, 5x at Anthropic, 6x at OpenAI, and up to 8x at Google — so generation-heavy workloads dominate real-world bills.
  • Five providers, one ladder: OpenAI, Anthropic, Google, xAI, and DeepSeek each ladder their lineups from premium reasoning tiers down to budget models, giving buyers substitutes at nearly every price point.

? The numbers:

  • GPT-5.5 (OpenAI) combined: $35.00
  • Claude Opus 4.8 (Anthropic) combined: $30.00
  • GPT-5.4 (OpenAI) combined: $17.50
  • Gemini 2.5 Pro (Google) combined: $11.25
  • Grok 4.3 (xAI) combined: $3.75
  • DeepSeek V4 Flash combined: $0.42
  • Flagship-to-floor spread: ~83x

Token pricing is now a strategic decision, not a line item. With a roughly 83x spread between GPT-5.5 ($35) and DeepSeek V4 Flash ($0.42), and output priced 2x to 8x above input depending on provider, routing the right task to the right model — rather than defaulting to the most capable one — is where AI teams find their largest and fastest cost savings.

Continue reading below for the full detailed article →

Overview

How 13 LLM APIs Stack Up on One Price Yardstick

API pricing has become the quiet battleground of the AI industry. Measured on one yardstick — the combined cost of 1M input plus 1M output tokens — 13 models from five providers stratify sharply: GPT-5.5 at $35 and Claude Opus 4.8 at $30 up top; workhorses like GPT-5.4 ($17.50) and Gemini 2.5 Pro ($11.25) in the middle; and a commodity floor held by Gemini 2.5 Flash-lite ($0.50) and DeepSeek V4 Flash ($0.42). That roughly 83x spread — with output priced 2x to 8x above input — means where you route a task matters more to your bill than almost any other engineering decision.

#1 OpenAI — $1.45 to $35 Across Four GPT Tiers

OpenAI fields the widest price ladder of the five providers: four GPT tiers spanning $35 down to $1.45 per combined 1M input + 1M output tokens — a 24x range inside a single API. At the top sits GPT-5.5, the most expensive model in this comparison and the rate every other API is priced against.

GPT-5.5 · $5 In / $30 Out per 1M Tokens

At $35 combined, GPT-5.5 is the most expensive API in this comparison, and the $30 output rate is what makes it so — six times its $5 input price. An agent that reads 1M tokens of context and writes 1M tokens of code pays 86% of its bill on the output side. The economics work for high-value reasoning — multi-step agents, hard debugging, architectural analysis — and stop working the moment the task becomes routine.

GPT-5.4 · $2.50 In / $15 Out per 1M Tokens

GPT-5.4 lands at $17.50 combined — exactly half of GPT-5.5 on both input and output, keeping the same 6x output premium. This is the workhorse tier: production chat, RAG answers, and everyday code generation where flagship reasoning is overkill. Teams that default here instead of GPT-5.5 cut token spend 50% across the board without changing a line of integration code.

GPT-5.4-mini · $0.75 In / $4.50 Out per 1M Tokens

At $5.25 combined, GPT-5.4-mini costs 30% of GPT-5.4 and just 15% of GPT-5.5. The tier is built for volume: summarization pipelines, draft generation, and tool-calling loops that fire thousands of requests a day. It also slides $0.75 under Anthropic's light-tier Claude Haiku 4.5 ($6 combined) — a mini priced to compete one class up.

GPT-5.4-nano · $0.20 In / $1.25 Out per 1M Tokens

GPT-5.4-nano closes OpenAI's ladder at $1.45 combined — cheap enough that classification, entity extraction, and request routing become near-free line items. It is not the market floor, though: DeepSeek V4 Flash ($0.42) and Gemini 2.5 Flash-lite ($0.50) both undercut it by roughly two-thirds. Nano's case is staying inside the OpenAI stack while pushing bulk workloads down to a $0.20 input rate.

GPT-5.5 by the Numbers

  • Input price per 1M tokens $5 What GPT-5.5 charges to read 1M tokens of prompt and context — tied exactly with Claude Opus 4.8's input rate.
  • Output price per 1M tokens $30 Six times the input rate — generation-heavy workloads like agents and code pay most of their GPT-5.5 bill on this line.
  • Combined, 1M in + 1M out $35 The highest combined rate of all 13 APIs — roughly 83x the $0.42 floor set by DeepSeek V4 Flash.

Frontier Lab

#2 Anthropic — Claude's Three Rungs, $6 to $30

Anthropic prices Claude on a strict three-rung ladder: Opus 4.8 at $30 combined, Sonnet 5 at $12, Haiku 4.5 at $6. It is the most consistent lineup in the dataset — output costs exactly 5x input at every tier ($5/$25, $2/$10, $1/$5) — which makes cost forecasting unusually predictable: whatever a task's input/output mix, moving one rung down cuts the bill by half or more.

Claude Opus 4.8 · $5 In / $25 Out per 1M Tokens

Opus 4.8 mirrors GPT-5.5's $5 input rate but charges $25 for output instead of $30, landing at $30 combined. That $5-per-million output gap compounds fast: a service generating 1B output tokens a month saves $5,000 monthly on Opus at identical volume. For agentic coding and long-form reasoning — the workloads flagship pricing is aimed at — that makes Opus the cheaper of the two frontier options on every generation-heavy job.

Claude Sonnet 5 · $2 In / $10 Out per 1M Tokens

Sonnet 5's $12 combined undercuts its closest OpenAI rival, GPT-5.4 ($17.50), by $5.50 — about 31% — while keeping input at just $2 per 1M tokens. It is the tier most coding assistants and production chat products actually run on: capable enough for multi-file code edits, priced low enough to serve as a default rather than an escalation.

Claude Haiku 4.5 · $1 In / $5 Out per 1M Tokens

Haiku 4.5 closes the ladder at $6 combined. Against other providers' light tiers it reads mid-priced — above GPT-5.4-mini ($5.25) and well above Gemini 2.5 Flash ($2.80) — but it keeps the same 5x output ratio as its Opus and Sonnet siblings. Teams already on Claude use it for triage, moderation, and short-answer endpoints where a $1 input rate keeps high-frequency calls affordable.

Big Tech Lab

#3 Google — Gemini From $11.25 Down to $0.50

Google prices Gemini a full class below the other US flagships: top-tier 2.5 Pro costs $11.25 combined — roughly a third of Claude Opus 4.8 — while Flash-lite bottoms out at $0.50. The trade hiding in the lineup is the output ratio: Gemini output runs 4x to more than 8x its input price, so the cheap headline rates favor read-heavy workloads over write-heavy ones.

Gemini 2.5 Pro · $1.25 In / $10 Out per 1M Tokens

At $11.25 combined, Gemini 2.5 Pro is the cheapest flagship in the comparison — but its 8x output-to-input ratio is the highest of any frontier model here. The pricing rewards a specific shape of work: long-context reading, document analysis, and RAG over large corpora, where inputs dwarf outputs and the $1.25 input rate does the heavy lifting. Flip the mix toward generation and the $10 output rate still undercuts Opus ($25) and GPT-5.5 ($30), but by less than the input numbers suggest.

Gemini 2.5 Flash · $0.30 In / $2.50 Out per 1M Tokens

Flash lands at $2.80 combined — 47% cheaper than GPT-5.4-mini ($5.25) and less than half of Claude Haiku 4.5 ($6). It shares a $2.50 output price with Grok 4.3 while charging under a quarter of Grok's $1.25 input rate. For high-throughput production endpoints — chat, summarization, structured extraction — Flash is the price-performance anchor most teams benchmark everything else against.

Gemini 2.5 Flash-lite · $0.10 In / $0.40 Out per 1M Tokens

Flash-lite's $0.10 input rate is the lowest of any model in the dataset, and its $0.50 combined cost is beaten only by DeepSeek V4 Flash ($0.42). At these prices token math nearly disappears from the meeting agenda: a billion input tokens costs $100. That makes it the default candidate for firehose workloads — log triage, spam filtering, bulk labeling — where per-call quality demands are modest but volume is enormous.

Challenger Lab

#4 xAI — Grok 4.3, One Model at $3.75 Combined

xAI is the only provider in this comparison listing a single price point: Grok 4.3 at $1.25 input and $2.50 output per 1M tokens. There is no flagship-to-nano ladder to route across — which simplifies integration to one decision, but leaves developers without an in-house budget fallback when volume spikes.

Grok 4.3 · $1.25 In / $2.50 Out per 1M Tokens

Grok 4.3's $3.75 combined places it between the mini and mid tiers — above Gemini 2.5 Flash ($2.80), below GPT-5.4-mini ($5.25). The standout number is its output multiple: at 2x input, it carries the lowest output premium of any US-provider model here, against 5x at Anthropic, 6x at OpenAI, and up to 8x at Google. For generation-heavy jobs — long answers, content drafting, verbose code — that ratio makes Grok cheaper in practice than its combined rank suggests.

It also matches Gemini 2.5 Pro's $1.25 input rate exactly while charging a quarter of its $10 output price. The catch is scope: with a single listing, teams that outgrow Grok's quality ceiling or need a sub-dollar tier must cross providers rather than move down a rung — a real integration cost that the per-token price doesn't capture.

Open-Weight Lab

#5 DeepSeek — The $0.42 Floor of LLM Pricing

DeepSeek's two listings — V4 Pro at $1.305 combined and V4 Flash at $0.42 — sit at the very bottom of the ladder. Both hold output at exactly 2x input, and even the Pro tier undercuts OpenAI's cheapest model, GPT-5.4-nano ($1.45). This is where the roughly 83x top-to-bottom spread of the whole comparison is anchored.

DeepSeek V4 Pro · $0.435 In / $0.87 Out per 1M Tokens

V4 Pro prices its premium tier at $1.305 combined — less than 4% of GPT-5.5's $35, and cheaper than every other provider's budget model except Gemini 2.5 Flash-lite ($0.50). For teams whose evaluations show acceptable quality on their task, the arithmetic is blunt: a workload costing $10,000 a month on GPT-5.4 reruns here for roughly $750 at the same token volume — with the usual caveats around latency, hosting jurisdiction, and compliance review.

DeepSeek V4 Flash · $0.14 In / $0.28 Out per 1M Tokens

V4 Flash defines the floor: $0.42 combined, roughly 1/83rd of GPT-5.5. A pipeline pushing 1M input and 1M output tokens a day costs $12.60 a month here versus $1,050 on GPT-5.5. At that spread the question stops being which model is best and becomes which tasks genuinely need frontier reasoning — because everything that doesn't can run at prices that barely register on a cloud bill.

Complete Pricing Data

All 13 LLM APIs, Ranked by Combined Cost

Provider, model, input price, output price, and combined cost for 1M input + 1M output tokens — the ranking metric used throughout this story. Reading the input and output columns side by side shows each provider's output premium: 2x at DeepSeek and xAI, 5x at Anthropic, 6x at OpenAI, and up to 8x at Google.

OpenaiGpt-5.553035
OpenaiGpt-5.42.51517.5
OpenaiGpt-5.4-mini0.754.55.25
OpenaiGpt-5.4-nano0.21.251.45
AnthropicClaude Opus 4.852530
AnthropicClaude Sonnet 521012
AnthropicClaude Haiku 4.5156
GoogleGemini 2.5 Pro1.251011.25
GoogleGemini 2.5 Flash0.32.52.8
GoogleGemini 2.5 Flash-lite0.10.40.5
XaiGrok 4.31.252.53.75
DeepseekDeepseek V4 Flash0.140.280.42
DeepseekDeepseek V4 Pro0.4350.871.305

Combined cost = input price per 1M tokens + output price per 1M tokens. Actual workload costs depend on your input/output token ratio; list prices exclude volume discounts, caching discounts, and batch pricing.

Why It Matters

Why LLM Pricing Matters Now

Token costs are the primary variable expense for AI products. For anything built on hosted LLM APIs, the per-token rate multiplied by monthly volume defines gross margin — and with rates spanning $0.42 to $35 per combined 1M tokens, two teams shipping the same feature can carry cost structures 83x apart. That is why pricing tables are now required reading for founders, CFOs, and engineering leads alike.

Who Pays per Token · Builders, Not Subscribers

Consumer subscriptions like a $20-a-month chatbot plan hide token economics behind a flat fee. API customers get no such buffer: every startup embedding an LLM in its product, every enterprise team running document pipelines, and every agent framework calling a model in a loop pays for exactly the tokens it moves. For these builders — SaaS products with AI features, coding assistants, support bots, internal automation — the pricing table is the cost of goods sold.

That is also why per-token rates matter more than subscription prices for this audience: a subscription caps one user's cost, while an API bill scales with every customer, every request, and every retry. A support bot answering 100,000 tickets a month at roughly 2,000 tokens each moves 200M tokens — about $1,750 on GPT-5.4 versus $280 on Gemini 2.5 Flash for the same traffic.

Where Token Volume Explodes · Agents, RAG, Pipelines

The workloads growing fastest are also the hungriest. Agent loops re-read their context on every step, RAG systems stuff retrieved documents into each prompt, and batch pipelines process entire archives — token volume compounds in ways a human chat session never does. A single agentic coding task can burn several hundred thousand tokens before it finishes, so the difference between a 2x and an 8x output premium starts deciding the architecture, not just the invoice.

The upside is that the market now offers substitutes at every tier. Five providers field 13 models across flagship, mid, mini, and nano classes, so buyers can arbitrage capability against cost — GPT-5.4-mini against Gemini 2.5 Flash, Claude Sonnet 5 against GPT-5.4 — in ways that were impossible when only one or two frontier models existed.

How to Read the Numbers · One Combined Metric

Providers quote input and output prices separately, which hides true cost — especially with output premiums ranging from 2x at DeepSeek and xAI to 8x at Google. Combining 1M input + 1M output tokens into one number levels the field for ranking, even if individual workloads skew toward one side. Read-heavy systems like RAG and long-document summarization should weight the input column; generation-heavy systems like code and drafting should weight output — and the ranking can flip between the two.

Conclusion

The Takeaway — Route Tasks, Don't Pick a Favorite

LLM API pricing now spans roughly 83x from top to bottom — $35 for GPT-5.5 down to $0.42 for DeepSeek V4 Flash per combined 1M input + 1M output tokens — making model selection the single largest cost lever in AI development.

The flagship race is tight: Claude Opus 4.8 ($30 combined) trails GPT-5.5 ($35) by $5 with identical $5 input pricing, so the real differentiation at the top happens on output rates and capability, not input cost.

Output premiums are a provider signature, not a constant: 2x input at DeepSeek and xAI, 5x at Anthropic, 6x at OpenAI, and up to 8x at Google. Real-world bills depend on how much a model writes, not just how much it reads — so match the ratio to your workload's shape.

For buyers, the winning strategy is routing, not loyalty: send each task to the cheapest tier that clears its quality bar, and revisit the table often — with five providers competing across 13 models, these numbers move fast.

Data Source and Attribution

The pricing data behind this story was compiled manually by FactsFigs from the official websites of the five providers covered — the public pricing pages of OpenAI, Anthropic, Google, xAI, and DeepSeek — with AI tools assisting the collection and cross-checking of per-model input and output rates. Because the dataset was assembled in-house rather than drawn from a single external source, no third-party source links apply. List prices change frequently — always verify current rates on each provider's official pricing page before committing spend.

FactsFigs reviews, cleans, and cross-checks every source dataset before shaping it into a data story. Each visualization is created and designed in FactsFigs Design Studio — an internal tool developed and owned by FactsFigs — and is the original work of a FactsFigs author, not an AI-generated copy of any existing graphic. Individual assets within a visual may or may not be produced with AI tools, but the design of the visual itself is solely FactsFigs' own.

Figures are estimates at the time of publication, provided for information only — nothing here is financial advice or a guarantee of accuracy.