Top AI Companies By Intelligence

By Saif Ur Rahman Published 22 Jun 2026 Updated 17 Jul 2026
FactsFigs Data Story
Top 10 AI Companies Ranked by FactsFigs Intelligence Score
FactsFigs

Facts & Figures

Visual Intelligence

TL;DR — Which AI Company Scores Highest for Intelligence

FactsFigs' private score, not a public benchmark

OpenAI's GPT-5.2 (xhigh) tops FactsFigs' internal intelligence ranking at 51, tested against identical prompts alongside nine other AI creators' flagship models. Anthropic (49) and Google (48) trail by only one to three points — the tightest three-way race in this ranking. Model count varies independently of peak score: Google tracks the most models (14) but ranks third, while KwaiKAT scores 36 with a single model.

! How to Read This Ranking:

  • Context: FactsFigs tested each creator's flagship model against the same set of prompts and scored the outputs privately — this is not MMLU, LMSYS, or any other public leaderboard.
  • Signal: The top three creators (OpenAI 51, Anthropic 49, Google 48) are separated by only three points, while scores fall steadily to KwaiKAT's 36 at #10.
  • Action: Read Best Intelligence and Models Count together — a high score signals peak flagship capability, while model count signals portfolio breadth, and the two don't always move together.

? Key Data at a Glance:

  • Top Score: 51 — OpenAI's GPT-5.2 (xhigh)
  • Creators Ranked: 10
  • Models Count Range: 1 (KwaiKAT) to 14 (Google)

This is FactsFigs' own private intelligence score, built from identical-prompt testing rather than a public benchmark suite — useful for comparing flagship models head-to-head, but not a substitute for standardized leaderboards like MMLU or LMSYS Arena.

Continue reading below for the full detailed article →

Overview

A Private Intelligence Score, Not a Public Benchmark

FactsFigs built this ranking by testing each of ten AI creators' flagship models against an identical set of prompts, then scoring the outputs in-house — a private evaluation, not a public benchmark suite like MMLU or LMSYS Chatbot Arena. OpenAI's GPT-5.2 (xhigh) topped that testing at 51, with Anthropic's Claude Opus 4.5 (49) and Google's Gemini 3 Pro Preview (48) close enough behind to call it a three-way race for the top spot. From there, scores fall in a fairly even line down to KwaiKAT's KAT-Coder-Pro V1 at 36, a 15-point spread across the full top ten. Models Count sits alongside each score for context — how many models a creator has in this dataset — and the two numbers tell different stories: a high score marks peak flagship capability, while a large model count marks how broadly a company productizes AI, from lightweight, low-cost variants up to the reasoning-heavy configuration that earns the headline score.

#1 OpenAI — Intelligence Score 51 Across 13 Models

OpenAI tops FactsFigs' private intelligence scoring at 51 — a score built from identical-prompt testing across every creator's flagship model, not a result pulled from a public leaderboard like MMLU or LMSYS. The lead over Anthropic (49) and Google (48) is narrow, but GPT-5.2 (xhigh) is the highest scorer in this internally compiled ranking, backed by the second-widest model portfolio in the field, trailing only Google.

Peak Score · GPT-5.2 (xhigh) at 51/100

51

GPT-5.2 (xhigh) is OpenAI's extended-reasoning configuration, and it posted the highest score of any model FactsFigs tested with identical prompts across all ten creators — a private benchmark, not a public leaderboard result.

13 Models Tracked · Second-Widest Lineup

13

OpenAI's 13 tracked models trail only Google's 14 in this dataset, spanning lighter, faster configurations up to the top-scoring xhigh reasoning tier evaluated here.

3-Point Lead · 51 vs. Anthropic's 49

3 pts

The gap over second-place Anthropic (49) and third-place Google (48) is narrow enough that a single new model release from either rival could flip the top spot in FactsFigs' next testing pass.

OpenAI by the Numbers

  • Score Ceiling · 51/100 51 A score of 51 sets the ceiling in FactsFigs' private testing pass, making OpenAI's GPT-5.2 (xhigh) the reference point every other creator's flagship is measured against.
  • Top-3 Spread · Just 3 Points 3-point spread OpenAI (51), Anthropic (49), and Google (48) are separated by only three points — the tightest three-way race in this ranking.
  • Widest Lineup · Google's 14 Models 14 Google lists more models than any other creator in this dataset — 14 — one ahead of OpenAI's 13, though a bigger lineup didn't translate into the top score.

Frontier Lab

#2 Anthropic — Claude Opus 4.5 Scores 49 From 6 Models

Anthropic keeps a lean lineup — 6 models in this dataset, less than half of OpenAI's 13 — and Claude Opus 4.5 is the one that carries the company to #2, scoring 49 in FactsFigs' identical-prompt testing, just two points off the top. Anthropic's models have built their reputation on careful, well-reasoned answers and strong coding performance rather than the widest possible feature set, and Opus 4.5's score reflects that focus: consistent output quality across long, multi-step prompts, the kind of task where staying coherent over many turns matters more than raw speed.

Multimodal AI

#3 Google — Gemini 3 Pro Scores 48 From the Widest Lineup

Google tracks more models than any other creator in this dataset — 14 — yet its best score, 48 from Gemini 3 Pro Preview (high), still lands third, one point behind Anthropic and three behind OpenAI. That gap between portfolio size and peak score fits Google's strategy: Gemini ships across Search, Workspace, and Android as much as it competes as a standalone flagship, so the lineup spans from lightweight Flash variants built for speed up to the high-effort Pro configuration scored here. In FactsFigs' testing, Gemini's strongest showings came on multimodal prompts — reading images, tables, and mixed context — more than on pure text reasoning.

Open-Weight Challenger

#4 Z AI — GLM-4.7 Scores 42 From 5 Models

Z AI's GLM-4.7 is the highest-scoring model from a five-model lineup, landing at 42 — six points behind Google and fifteen off the top score. GLM models have built a name as capable open-weight options, particularly for bilingual Chinese-English tasks and lower-cost deployment compared with closed frontier labs, and GLM-4.7's score in FactsFigs' testing held up on general reasoning prompts even without the multi-tier model ladder that OpenAI, Anthropic, and Google maintain.

Cost-Efficient Reasoning

#5 DeepSeek — V3.2 Ties for 41 From 8 Models

DeepSeek V3.2 scored 41 in FactsFigs' testing, tying xAI's Grok 4 for fifth place, out of an 8-model lineup — the third-widest in this ranking behind Google and OpenAI. DeepSeek built its reputation on open-weight models that approach frontier-lab reasoning at a fraction of the training and inference cost, and V3.2's score reflects that: it held its own on math and step-by-step reasoning prompts against labs running far larger infrastructure budgets.

Real-Time Reasoning

#6 xAI — Grok 4 Matches DeepSeek at 41

Grok 4 ties DeepSeek V3.2 exactly at 41, drawn from a 7-model lineup. xAI's models lean on real-time access to X (formerly Twitter) data and a more conversational, less-filtered response style rather than chasing the widest benchmark suite, and that shows up in FactsFigs' testing as a relative strength on current-events and reasoning-under-ambiguity prompts — the kind of question where up-to-the-minute context matters more than a structured coding task.

Agentic Tool Use

#7 Kimi — K2 Thinking Scores 40 From Just 3 Models

Moonshot AI's Kimi fields the leanest lineup of the top seven — just 3 models — with Kimi K2 Thinking topping the group at 40. Kimi's models are built around agentic tool use: chaining searches, running code, and planning multi-step tasks inside a single response rather than answering in one shot, and that specialization is exactly where K2 Thinking held its ground against flagship models from labs with far larger model portfolios.

Efficient Agents

#8 MiniMax — M2.1 Scores 39 From Only 2 Models

MiniMax fields just two models in this dataset, and MiniMax-M2.1 is the stronger of the pair, scoring 39 — tied with Xiaomi for the eighth and ninth spots. MiniMax has positioned its models as efficient agentic workers built for fast tool-calling and structured task execution rather than the deepest possible reasoning chains, trading some peak score for lower latency and cost per task.

Lightweight, Mobile-First

#9 Xiaomi — MiMo-V2-Flash Ties MiniMax at 39

Xiaomi's MiMo-V2-Flash ties MiniMax-M2.1 at 39, also from a two-model lineup. True to the "Flash" naming, Xiaomi's models are tuned for speed and on-device efficiency rather than maximum reasoning depth — a fit for Xiaomi's hardware business, where models need to run well on phones and lightweight edge deployments rather than top a leaderboard.

Coding Specialist

#10 KwaiKAT — KAT-Coder-Pro V1 Scores 36 as a Single Model

KwaiKAT closes the top ten with a single model in this dataset — KAT-Coder-Pro V1 — scoring 36. The name gives away the focus: it's built specifically for coding tasks rather than general-purpose use, and that narrow, single-model approach explains both its lower peak score against nine generalist rivals and its sharper edge on code completion and debugging prompts, where broad conversational ability matters less than getting the syntax right.

Full Rankings

All 10 AI Companies, Ranked by FactsFigs Intelligence Score

Rank, creator, best-scoring model, FactsFigs Intelligence Score, and total models tracked for each of the ten AI companies in this private ranking.

1OpenAIGPT-5.2 (xhigh)5113
2AnthropicClaude Opus 4.5496
3GoogleGemini 3 Pro Preview (high)4814
4Z AIGLM-4.7425
5DeepSeekDeepSeek V3.2418
6xAIGrok 4417
7KimiKimi K2 Thinking403
8MiniMaxMiniMax-M2.1392
9XiaomiMiMo-V2-Flash392
10KwaiKATKAT-Coder-Pro V1361

Scores and model counts come from FactsFigs' own internal, identical-prompt testing, not a public benchmark suite — 'Best Intelligence' is each creator's peak score in this evaluation, and 'Models Count' measures portfolio breadth only.

Takeaway

The Takeaway — A Private Score, Not a Standard Benchmark

FactsFigs' intelligence score is not MMLU, not LMSYS Arena, and not any other public leaderboard — it's a private ranking built from identical-prompt testing across ten AI creators' flagship models, scored in-house. OpenAI's GPT-5.2 (xhigh) topped that testing at 51, with Anthropic's Claude Opus 4.5 (49) and Google's Gemini 3 Pro Preview (48) close enough behind that the order could shift with the next model release from either rival.

Model count tells a separate story from peak score. Google tracks the most models in this dataset (14) and OpenAI the second-most (13), yet neither leads on raw intelligence — KwaiKAT closes the ranking with a single model and still lands a respectable 36. A large model lineup signals product breadth, not necessarily a sharper flagship, and reading the two numbers together is closer to the real picture than either alone.

Data Source and Attribution

The intelligence scores behind this story are FactsFigs' own private ranking, not a public benchmark suite like MMLU or LMSYS Chatbot Arena. Each of the ten creators' flagship models was tested against the same set of prompts, and the resulting output quality was scored and compiled in-house by a FactsFigs evaluator. Because the ranking was assembled internally rather than pulled from an existing benchmark, no third-party dataset link applies.

FactsFigs reviews, cleans, and cross-checks every source dataset before shaping it into a data story. Each visualization is created and designed in FactsFigs Design Studio — an internal tool developed and owned by FactsFigs — and is the original work of a FactsFigs author, not an AI-generated copy of any existing graphic. Individual assets within a visual may or may not be produced with AI tools, but the design of the visual itself is solely FactsFigs' own.

Figures are estimates at the time of publication, provided for information only — nothing here is financial advice or a guarantee of accuracy.

2026-07-17