Services
Contact Us

LLM Benchmarks

One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations.

Last Updated Jun 2026
TOP MODEL
Claude Fable 5
Index 92
Best value
MiniMax M3
$0.42 / 1M
Fastest
GPT OSS 120B
0.19s TTFT
Coverage
64 models
8 benchmarks · 336 eval runs
AIMultiple Intelligence Index

Leaderboard

The highest-scoring models across all benchmarks.

Filter & Sort
#
Model
Index
LegalBench
FinanceReasoning
Text-to-SQL
Agentic RAG
ARC-AGI-2
Agentic LLM Benchmark
FrontierMath
Swe-Bench
1
Claude Fable 5
Claude Fable 5
Anthropic
92
89909098-69--
2
Kimi K3
Kimi K3
Moonshot AI
88
86887793-73--
3
GPT-5.6 Sol
GPT-5.6 Sol
OpenAI
87
879074879362--
4
Grok 4.5
Grok 4.5
X
86
86887983-73--
5
GPT-5.5
GPT-5.5
OpenAI
83
87---8559--
6
GPT-5.6 Terra
GPT-5.6 Terra
OpenAI
80
858771918461--
7
GPT-5.6 Sol Pro
GPT-5.6 Sol Pro
OpenAI
80
-917996-54--
8
Claude Opus 4.6
Claude Opus 4.6
Anthropic
78
-8868806972--
9
Gemini 3 Pro Preview
Gemini 3 Pro Preview
Google
77
8786608931---
10
Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview
Google
77
87876589774627-
Page 1 of 7

Cost$20.00
Latency3.98s
Context1M
TTFT3.98s
LegalBench
89
FinanceReasoning
90
Text-to-SQL
90
Agentic RAG
98
Agentic LLM Benchmark
69

Cost$6.00
Latency252.11s
Context1M
TTFT252.11s
LegalBench
86
FinanceReasoning
88
Text-to-SQL
77
Agentic RAG
93
Agentic LLM Benchmark
73

Cost$11.25
Latency2.47s
Context1M
TTFT2.47s
LegalBench
87
FinanceReasoning
90
Text-to-SQL
74
Agentic RAG
87
ARC-AGI-2
93
Agentic LLM Benchmark
62

Cost$3.00
Latency10.26s
Context500k
TTFT10.26s
LegalBench
86
FinanceReasoning
88
Text-to-SQL
79
Agentic RAG
83
Agentic LLM Benchmark
73

Cost$11.25
Latency1.03s
Context272k
TTFT1.03s
LegalBench
87
ARC-AGI-2
85
Agentic LLM Benchmark
59

Cost$5.63
Latency1.92s
Context1M
TTFT1.92s
LegalBench
85
FinanceReasoning
87
Text-to-SQL
71
Agentic RAG
91
ARC-AGI-2
84
Agentic LLM Benchmark
61

Cost$11.25
Latency11.46s
Context1M
TTFT11.46s
FinanceReasoning
91
Text-to-SQL
79
Agentic RAG
96
Agentic LLM Benchmark
54

Cost$10.00
Latency1.75s
Context1M
TTFT1.75s
FinanceReasoning
88
Text-to-SQL
68
Agentic RAG
80
ARC-AGI-2
69
Agentic LLM Benchmark
72

Cost$4.50
Latency-
Context1M
TTFT
LegalBench
87
FinanceReasoning
86
Text-to-SQL
60
Agentic RAG
89
ARC-AGI-2
31

Cost$4.50
Latency33.83s
Context1M
TTFT33.83s
LegalBench
87
FinanceReasoning
87
Text-to-SQL
65
Agentic RAG
89
ARC-AGI-2
77
Agentic LLM Benchmark
46
FrontierMath
27
Page 1 of 7

Charts

Following "Best Performing" models

Frontier Over Time
Intelligence Index by model release date
Cost vs Performance
Blended cost against Intelligence Index
Model × Benchmark
Top models on selected benchmarks
Methodology

How The Index is Built

The Intelligence Index normalizes each benchmark to a 0-100 scale and averages a model's relative standing across the benchmarks it was evaluated on. Each benchmark below feeds that score.

Recent Updates

Recent Updates

Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.

Moonshot AI

Kimi K3

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-5.6 Sol

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-5.6 Terra

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-5.6 Sol Pro

New model added to the AIMultiple Intelligence Index.

Explore LLM Use Cases, Analyses & Benchmarks

Audience Simulation: Can LLMs Predict Human Behavior?

LLM
Benchmark
Jun 22

In marketing, evaluating how accurately LLMs predict human behavior is crucial for assessing their effectiveness in anticipating audience needs and recognizing the risks of misalignment, ineffective communication, or unintended influence. Audience simulation with LLMs enables the modeling of virtual audiences, helping organizations anticipate reactions to content or products without relying on costly surveys or focus…

Read More
LLM
Benchmark
Jun 15

AI Gateways for OpenAI: OpenRouter Alternatives

We benchmarked OpenRouter, SambaNova, TogetherAI, Groq, and AI/ML API across three indicators (first-token latency, total latency, and output-token count), with 300 tests using short prompts (approx. 18 tokens) and long prompts (approx. 203 tokens) for total latency. If you plan to use one of these AI gateways, you can: In this benchmark, we compared OpenRouter,…

LLM
Benchmark
Jun 5

Large Language Models in Cybersecurity

We evaluated 7 large language models across 9 cybersecurity domains using SecBench, a large-scale and multi-format benchmark for security tasks. We tested each model on 44,823 multiple-choice questions (MCQs) and 3,087 short-answer questions (SAQs), covering data security, identity & access management, network security, vulnerability management, and cloud security. MCQs (Multiple-Choice Questions) benchmarking: SAQs (Short Answer…

LLM
Insight
May 26

ChatGPT for Customer Service: Top 10 Use Cases

ChatGPT has moved from novelty to infrastructure in customer service. Companies are using it to cut response times, handle volume their teams can’t absorb, and reduce the cost of routine interactions. But results vary sharply depending on how it’s implemented. OpenAI launched GPT-5.2, a materially more capable model that is better at instruction-following, reasoning across…

LLM
Benchmark
Apr 15

LLM Quantization: BF16 vs FP8 vs INT4

We benchmarked Qwen3-32B at 4 precision levels (BF16, FP8, GPTQ-Int8, GPTQ-Int4) on a single NVIDIA H100 80GB GPU. Each configuration was evaluated on 2 benchmarks (~12.2K questions) covering knowledge and code generation, plus 2,000+ inference runs to measure throughput. Int4 is 2.7x faster than BF16 while losing less than 2 points on MMLU-Pro, but code…