Services
Contact Us

LLM Benchmarks

One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations. See the methodology.

TOP MODEL
Claude Opus 5
Index 86
Best value
DeepSeek V4 Flash 0731
$0.11 / 1M
Fastest
Trinity Large Thinking
0.12s TTFT
Coverage
151 models
12 benchmarks · 697 eval runs
AIMultiple Intelligence Index

Leaderboard

The highest-scoring models across all benchmarks.

Filter & Sort
#
Model
Index
Agentic RAG
ARC-AGI-1
ARC-AGI-2
ArxivMath
CritPt
FinanceReasoning
FrontierMath
LegalBench
MMMU-Pro
Terminal-Bench 2.1
Terminal-Bench Hard
Text-to-SQL
1
Claude Opus 5
Claude Opus 5
Anthropic
86
-989091299073878589--
2
GPT-5.6 Sol
GPT-5.6 Sol
OpenAI
81
879893-329083878389-74
3
Claude Fable 5
Claude Fable 5
Anthropic
79
989989792990888981856390
4
GPT-5.6 Terra
GPT-5.6 Terra
OpenAI
67
919784-308771858187-71
5
Claude Opus 4.8
Claude Opus 4.8
Anthropic
66
100937271-89568479855877
6
Claude Fable 5.1
Claude Fable 5.1
Anthropic
66
-9890-31-8889-91--
7
Gemini 3.7 Flash
Gemini 3.7 Flash
Google
62
-9685-14-37878686--
8
GPT-5.5 Pro
GPT-5.5 Pro
OpenAI
62
-9785-31-78-----
9
Kimi K3
Kimi K3
Moonshot AI
61
939560-238839868285-77
10
GPT-5.6 Sol Pro
GPT-5.6 Sol Pro
OpenAI
60
96----9181----79
Page 1 of 16

Cost$10.00
Latency3.83s
Context1M
TTFT3.83s
ARC-AGI-1
98
ARC-AGI-2
90
ArxivMath
91
CritPt
29
FinanceReasoning
90
FrontierMath
73
LegalBench
87
MMMU-Pro
85
Terminal-Bench 2.1
89

Cost$8.00
Latency4.43s
Context1M
TTFT4.43s
Agentic RAG
87
ARC-AGI-1
98
ARC-AGI-2
93
CritPt
32
FinanceReasoning
90
FrontierMath
83
LegalBench
87
MMMU-Pro
83
Terminal-Bench 2.1
89
Text-to-SQL
74

Cost$20.00
Latency5.28s
Context1M
TTFT5.28s
Agentic RAG
98
ARC-AGI-1
99
ARC-AGI-2
89
ArxivMath
79
CritPt
29
FinanceReasoning
90
FrontierMath
88
LegalBench
89
MMMU-Pro
81
Terminal-Bench 2.1
85
Terminal-Bench Hard
63
Text-to-SQL
90

Cost$4.50
Latency1.92s
Context1M
TTFT1.92s
Agentic RAG
91
ARC-AGI-1
97
ARC-AGI-2
84
CritPt
30
FinanceReasoning
87
FrontierMath
71
LegalBench
85
MMMU-Pro
81
Terminal-Bench 2.1
87
Text-to-SQL
71

Cost$10.00
Latency4.53s
Context1M
TTFT4.53s
Agentic RAG
100
ARC-AGI-1
93
ARC-AGI-2
72
ArxivMath
71
FinanceReasoning
89
FrontierMath
56
LegalBench
84
MMMU-Pro
79
Terminal-Bench 2.1
85
Terminal-Bench Hard
58
Text-to-SQL
77

Cost$20.00
Latency5.44s
Context1M
TTFT5.44s
ARC-AGI-1
98
ARC-AGI-2
90
CritPt
31
FrontierMath
88
LegalBench
89
Terminal-Bench 2.1
91

Cost$1.50
Latency6.05s
Context1M
TTFT6.05s
ARC-AGI-1
96
ARC-AGI-2
85
CritPt
14
FrontierMath
37
LegalBench
87
MMMU-Pro
86
Terminal-Bench 2.1
86

Cost$33.75
Latency3.65s
Context1M
TTFT3.65s
ARC-AGI-1
97
ARC-AGI-2
85
CritPt
31
FrontierMath
78

Cost$5.44
Latency0.93s
Context1M
TTFT0.93s
Agentic RAG
93
ARC-AGI-1
95
ARC-AGI-2
60
CritPt
23
FinanceReasoning
88
FrontierMath
39
LegalBench
86
MMMU-Pro
82
Terminal-Bench 2.1
85
Text-to-SQL
77

Cost$2.00
Latency10.25s
Context1M
TTFT10.25s
Agentic RAG
96
FinanceReasoning
91
FrontierMath
81
Text-to-SQL
79
Page 1 of 16
Frontier Over Time
Intelligence Index by model release date
Cost vs Performance
Blended cost against Intelligence Index
Model × Benchmark
Top models on selected benchmarks
Methodology

How The Index is Built

Each benchmark becomes a 0-100 standing, and the index is the weighted average of those standings. A benchmark a model has not been run on counts as the field average, so scores stay comparable. Index = Terminal-Bench-Science (14%) + FrontierMath (12%) + Frontier-Bench (12%) + ArxivMath (10%) + MMMU-Pro (9%) + ARC-AGI-2 (8%) + CritPt (6%) + FinanceReasoning (5%) + Text-to-SQL (5%) + BioMysteryBench (5%) + ARC-AGI-1 (5%) + Terminal-Bench 2.1 (3%) + Agentic RAG (3%) + Terminal-Bench Hard (2%) + LegalBench (1%).

Recent Updates

Recent Updates

Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.

Google

Gemini 3.8 Flash

New model added to the AIMultiple Intelligence Index.

Anthropic

Claude Fable 5.1

New model added to the AIMultiple Intelligence Index.

Tencent

Hy4 preview

New model added to the AIMultiple Intelligence Index.

Z AI

GLM 5.3 Flash

New model added to the AIMultiple Intelligence Index.

AI Model Release TimelineLast 12 models in 23 days

Key AI model releases from leading providers over the past 23 days.

Provider
OpenAI
OpenAI
04 Sep 2026
GPT-6 Astra +1 moreGPT-6 AstraGPT-6 Astra Pro
Google
Google
02 Sep 2026
Gemini 3.8 Flash
Meta
Meta
21 Aug 2026
Muse Spark 1.2 Contributor
02 Sep 2026
Muse Spark 1.3 +1 moreMuse Spark 1.3Muse Spark 1.3 Contributor
Anthropic
Anthropic
01 Sep 2026
Claude Fable 5.1
Tencent
Tencent
28 Aug 2026
Hy4 preview
Alibaba Cloud
Alibaba Cloud
14 Aug 2026
Qwen3.8 27B
26 Aug 2026
Qwen3.8 Flash
Z AI
Z AI
18 Aug 2026
GLM 5.3
26 Aug 2026
GLM 5.3 Flash

Explore LLM Use Cases, Analyses & Benchmarks

Benchmark of 40+ LLMs in Finance: Claude Fable 5 & GPT-5.6 Sol

LLM
Benchmark
Aug 14

We evaluated LLMs on 238 hard questions from the FinanceReasoning benchmark (Tang et al.).1 This subset targets the most challenging financial-reasoning tasks, assessing complex, multi-step quantitative reasoning involving financial concepts and formulas. Our evaluation employed a custom prompt design and scoring criteria of accuracy and token consumption. For a detailed explanation of how these metrics…

Read More
LLM
Benchmark
Aug 12

LLM Latency Benchmark by Use Cases

We benchmarked 11 top large language models with a total of 1,320 requests, splitting reasoning and non-reasoning models, and measured first-token latency, per-token latency, and overall response time. You can find details on how we measured latency here. We report reasoning and non-reasoning models separately. Reasoning models spend several seconds thinking before the first visible…

LLM
Benchmark
Aug 11

Agentic IT: Can LLMs Design a Benchmark

We gave 12 large language models the job a benchmark team does: invent a benchmark, build it, run four models through it, and report the results. Each did it twice. None of the 24 attempts passed every criterion, and six of the rubric’s checks were passed by none of them. The two topics are text-to-SQL,…

LLM
Benchmark
Aug 4

Compare Multimodal AI Models on Visual Reasoning

We benchmarked 15 leading multimodal AI models on visual reasoning using 200 visual-based questions. The evaluation consisted of two tracks: 100 chart understanding questions testing data visualization interpretation, and 100 visual logic questions assessing pattern recognition and spatial reasoning. Each question was run 5 times to ensure consistent and reliable results. See our benchmark methodology…

LLM
Benchmark
Aug 2

AI Gateways for OpenAI: OpenRouter Alternatives

We benchmarked OpenRouter, SambaNova, TogetherAI, Groq, and AI/ML API across three indicators (first-token latency, total latency, and output-token count), with 300 tests using short prompts (approx. 18 tokens) and long prompts (approx. 203 tokens) for total latency. If you plan to use one of these AI gateways, you can: In this benchmark, we compared OpenRouter,…

LLM
Benchmark
Jun 5

Large Language Models in Cybersecurity

We evaluated 7 large language models across 9 cybersecurity domains using SecBench, a large-scale and multi-format benchmark for security tasks. We tested each model on 44,823 multiple-choice questions (MCQs) and 3,087 short-answer questions (SAQs), covering data security, identity & access management, network security, vulnerability management, and cloud security. MCQs (Multiple-Choice Questions) benchmarking: SAQs (Short Answer…

LLM
Benchmark
Apr 15

LLM Quantization: BF16 vs FP8 vs INT4

We benchmarked Qwen3-32B at 4 precision levels (BF16, FP8, GPTQ-Int8, GPTQ-Int4) on a single NVIDIA H100 80GB GPU. Each configuration was evaluated on 2 benchmarks (~12.2K questions) covering knowledge and code generation, plus 2,000+ inference runs to measure throughput. Int4 is 2.7x faster than BF16 while losing less than 2 points on MMLU-Pro, but code…