Premium
Services
Premium

LLM Benchmarks

One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations. See the methodology.

TOP MODEL
Claude Opus 5
Index 86
Best value
DeepSeek V4 Flash 0731
$0.11 / 1M
Fastest
Trinity Large Thinking
0.12s TTFT
Coverage
151 models
12 benchmarks · 697 eval runs
AIMultiple Intelligence Index

Leaderboard

The highest-scoring models across all benchmarks.

Filter & Sort
#
Model
Index
Agentic RAG
ARC-AGI-1
ARC-AGI-2
ArxivMath
CritPt
FinanceReasoning
FrontierMath
LegalBench
MMMU-Pro
Terminal-Bench 2.1
Terminal-Bench Hard
Text-to-SQL
1
Claude Opus 5
Claude Opus 5
Anthropic
86
-989091299073878589--
2
GPT-5.6 Sol
GPT-5.6 Sol
OpenAI
81
879893-329083878389-74
3
Claude Fable 5
Claude Fable 5
Anthropic
79
989989792990888981856390
4
GPT-5.6 Terra
GPT-5.6 Terra
OpenAI
67
919784-308771858187-71
5
Claude Opus 4.8
Claude Opus 4.8
Anthropic
66
100937271-89568479855877
6
Claude Fable 5.1
Claude Fable 5.1
Anthropic
66
-9890-31-8889-91--
7
Gemini 3.7 Flash
Gemini 3.7 Flash
Google
62
-9685-14-37878686--
8
GPT-5.5 Pro
GPT-5.5 Pro
OpenAI
62
-9785-31-78-----
9
Kimi K3
Kimi K3
Moonshot AI
61
939560-238839868285-77
10
GPT-5.6 Sol Pro
GPT-5.6 Sol Pro
OpenAI
60
96----9181----79
Page 1 of 16

Cost$10.00
Latency3.83s
Context1M
TTFT3.83s
ARC-AGI-1
98
ARC-AGI-2
90
ArxivMath
91
CritPt
29
FinanceReasoning
90
FrontierMath
73
LegalBench
87
MMMU-Pro
85
Terminal-Bench 2.1
89

Cost$8.00
Latency4.43s
Context1M
TTFT4.43s
Agentic RAG
87
ARC-AGI-1
98
ARC-AGI-2
93
CritPt
32
FinanceReasoning
90
FrontierMath
83
LegalBench
87
MMMU-Pro
83
Terminal-Bench 2.1
89
Text-to-SQL
74

Cost$20.00
Latency5.28s
Context1M
TTFT5.28s
Agentic RAG
98
ARC-AGI-1
99
ARC-AGI-2
89
ArxivMath
79
CritPt
29
FinanceReasoning
90
FrontierMath
88
LegalBench
89
MMMU-Pro
81
Terminal-Bench 2.1
85
Terminal-Bench Hard
63
Text-to-SQL
90

Cost$4.50
Latency1.92s
Context1M
TTFT1.92s
Agentic RAG
91
ARC-AGI-1
97
ARC-AGI-2
84
CritPt
30
FinanceReasoning
87
FrontierMath
71
LegalBench
85
MMMU-Pro
81
Terminal-Bench 2.1
87
Text-to-SQL
71

Cost$10.00
Latency4.53s
Context1M
TTFT4.53s
Agentic RAG
100
ARC-AGI-1
93
ARC-AGI-2
72
ArxivMath
71
FinanceReasoning
89
FrontierMath
56
LegalBench
84
MMMU-Pro
79
Terminal-Bench 2.1
85
Terminal-Bench Hard
58
Text-to-SQL
77

Cost$20.00
Latency5.44s
Context1M
TTFT5.44s
ARC-AGI-1
98
ARC-AGI-2
90
CritPt
31
FrontierMath
88
LegalBench
89
Terminal-Bench 2.1
91

Cost$1.50
Latency6.05s
Context1M
TTFT6.05s
ARC-AGI-1
96
ARC-AGI-2
85
CritPt
14
FrontierMath
37
LegalBench
87
MMMU-Pro
86
Terminal-Bench 2.1
86

Cost$33.75
Latency3.65s
Context1M
TTFT3.65s
ARC-AGI-1
97
ARC-AGI-2
85
CritPt
31
FrontierMath
78

Cost$5.44
Latency0.93s
Context1M
TTFT0.93s
Agentic RAG
93
ARC-AGI-1
95
ARC-AGI-2
60
CritPt
23
FinanceReasoning
88
FrontierMath
39
LegalBench
86
MMMU-Pro
82
Terminal-Bench 2.1
85
Text-to-SQL
77

Cost$2.00
Latency10.25s
Context1M
TTFT10.25s
Agentic RAG
96
FinanceReasoning
91
FrontierMath
81
Text-to-SQL
79
Page 1 of 16
Frontier Over Time
Intelligence Index by model release date
Cost vs Performance
Blended cost against Intelligence Index
Model × Benchmark
Top models on selected benchmarks
Methodology

How The Index is Built

Each benchmark becomes a 0-100 standing, and the index is the weighted average of those standings. A benchmark a model has not been run on counts as the field average, so scores stay comparable. Index = Terminal-Bench-Science (14%) + FrontierMath (12%) + Frontier-Bench (12%) + ArxivMath (10%) + MMMU-Pro (9%) + ARC-AGI-2 (8%) + CritPt (6%) + FinanceReasoning (5%) + Text-to-SQL (5%) + BioMysteryBench (5%) + ARC-AGI-1 (5%) + Terminal-Bench 2.1 (3%) + Agentic RAG (3%) + Terminal-Bench Hard (2%) + LegalBench (1%).

Recent Updates

Recent Updates

Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.

Google

Gemini 3.8 Flash

New model added to the AIMultiple Intelligence Index.

Anthropic

Claude Fable 5.1

New model added to the AIMultiple Intelligence Index.

Tencent

Hy4 preview

New model added to the AIMultiple Intelligence Index.

Z AI

GLM 5.3 Flash

New model added to the AIMultiple Intelligence Index.

AI Model Release TimelineLast 12 models in 21 days

Key AI model releases from leading providers over the past 21 days.

Provider
OpenAI
OpenAI
04 Sep 2026
GPT-6 Astra +1 moreGPT-6 AstraGPT-6 Astra Pro
Alibaba Cloud
Alibaba Cloud
26 Aug 2026
Qwen3.8 Flash
03 Sep 2026
Qwen3.8 Max (0902)
Google
Google
02 Sep 2026
Gemini 3.8 Flash
Meta
Meta
21 Aug 2026
Muse Spark 1.2 Contributor
02 Sep 2026
Muse Spark 1.3 +1 moreMuse Spark 1.3Muse Spark 1.3 Contributor
Anthropic
Anthropic
01 Sep 2026
Claude Fable 5.1
Tencent
Tencent
28 Aug 2026
Hy4 preview
Z AI
Z AI
18 Aug 2026
GLM 5.3
26 Aug 2026
GLM 5.3 Flash

Explore LLM Use Cases, Analyses & Benchmarks

LLM Automation: Top 7 Tools & 8 Case Studies 

LLM
Insight
Aug 27

LLM automation refers to shift to intelligent automation tools that leverage LLMs, including AI agents, fine-tuned LLMs and RAG models to automate and coordinate tasks. Explore what LLM automation is, its top real-life applications and major tools: Large language models in automation is a systematic approach that combines Natural Language Processing (NLP) with existing process…

Read More
LLM
Open World Evaluation
Aug 26

LLM Orchestration: 22 Frameworks and Gateways

Optimizing LLM orchestration is key to improving performance while keeping resource use under control. To evaluate how different orchestration approaches perform in practice, we benchmarked: Discover selected LLM orchestration tools, including developer frameworks and enterprise gateways: LLM Orchestration involves managing and integrating multiple Large Language Models (LLMs) to perform complex tasks efficiently. It ensures smooth…

LLM
Benchmark
Aug 24

AIM Enterprise: Agentic Enterprise Benchmark

Enterprises use LLMs every day for their regular tasks. To find the most cost-efficient LLMs, we designed AIM Enterprise, an agentic enterprise benchmark, where we used 69 real enterprise tasks across strategy, marketing, HR, sales, and operations. Two judge models scored every file, and they often disagreed. On about a third of the individual scores…

LLM
Benchmark
Aug 24

Text-to-SQL: Comparison of LLM Accuracy

We ran 36 large language models over 759 questions from BIRD-SQL, each model writing SQL against a database it had to identify for itself out of 11 candidates. Every parseable query was executed against the real database and its result set compared with the result set of BIRD’s gold query. Missing, malformed and execution-failing queries…

LLM
Insight
Aug 21

LLM Parameters: GPT-5 High, Medium, Low and Minimal

Some LLMs, such as OpenAI’s GPT-5 family, come in different versions (e.g., GPT-5, GPT-5-mini, and GPT-5-nano) and with various parameter settings, including high, medium, low, and minimal. Below, we explore the differences between these model versions by gathering their benchmark performance and the costs to run the benchmarks. We used the GPT-5 family in our…

LLM
Benchmark
Aug 21

HALC-Bench: LLM Hallucination on Long-Context Retrieval Benchmark

HALC-Bench (LLM Hallucination on Long-Context Retrieval Benchmark) measures a large language model’s resistance to fabricating evidence for a metric that does not exist in the target document by using 3 haystacks placed at the beginning, middle, and end of the model’s context window, with 204 questions. claude-fable-5 answered all 204 traps correctly at every haystack…

LLM
Feature Comparison
Aug 21

Top LLMOps Tools & Compare them to MLOPs

LLMOps platforms handle the operational side of running large language models: deployment, monitoring, evaluation, and cost management. We examined top LLMOps tools, their core features, pricing models, and how they differ from each other to help identify the best fit for various use cases. A breakdown of each metric is provided below: LLMOps platforms support…

LLM
Insight
Aug 19

50+ ChatGPT Use Cases with Real Life Examples

ChatGPT reached approximately 1 billion weekly active users in early 2026 roughly 10% of the world’s population.63 OpenAI surpassed $20 billion in annual revenue for 2025, confirmed by CFO Sarah Friar.66 The Anthropic Economic Index distinguishes two modes of use: augmentation, in which a human interacts with AI, and automation, in which AI completes tasks…

LLM
Insight
Aug 19

ChatGPT for Customer Service: Top 10 Use Cases

ChatGPT has moved from novelty to infrastructure in customer service. Companies are using it to cut response times, handle volume their teams can’t absorb, and reduce the cost of routine interactions. But results vary sharply depending on how it’s implemented. OpenAI launched GPT-5.6, a materially more capable model that is better at instruction-following, reasoning across…

LLM
Insight
Aug 18

The Future of Large Language Models

See the future of large language models by delving into promising approaches, such as self-training, fact-checking, and sparse expertise that could address LLM limitations. Success rate comparison of LLM’s Claude Sonnet 4.6 led the benchmark with an overall score of 0.748, with base and thinking variants tied to three decimal places. Claude Opus 4.8 (0.702),…

LLM
Benchmark
Aug 16

Intelligence Density of 71 LLMs for Smarter & Denser Models

We tracked 71 LLMs released between February 2023 and May 2026 and collected 10 public benchmarks to measure intelligence density. We divided the capability score by the resource the model consumes (active parameters, training compute, and inference price). To calculate intelligence density, we executed the following steps: See methodology for the scoring approach, and per-resource…