Services
Contact Us

LLM Benchmarks

One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations. See the methodology.

TOP MODEL
Claude Fable 5
Index 92
Best value
MiniMax M3
$0.42 / 1M
Fastest
GPT OSS 120B
0.23s TTFT
Coverage
64 models
8 benchmarks · 336 eval runs
AIMultiple Intelligence Index

Leaderboard

The highest-scoring models across all benchmarks.

Filter & Sort
#
Model
Index
LegalBench
FinanceReasoning
Text-to-SQL
Agentic RAG
ARC-AGI-2
Agentic LLM Benchmark
FrontierMath
Swe-Bench
1
Claude Fable 5
Claude Fable 5
Anthropic
92
89909098-69--
2
Kimi K3
Kimi K3
Moonshot AI
88
86887793-73--
3
GPT-5.6 Sol
GPT-5.6 Sol
OpenAI
87
879074879362--
4
Grok 4.5
Grok 4.5
X
86
86887983-73--
5
GPT-5.5
GPT-5.5
OpenAI
83
87---8559--
6
GPT-5.6 Terra
GPT-5.6 Terra
OpenAI
80
858771918461--
7
GPT-5.6 Sol Pro
GPT-5.6 Sol Pro
OpenAI
80
-917996-54--
8
Claude Opus 4.6
Claude Opus 4.6
Anthropic
78
-8868806972--
9
Gemini 3 Pro Preview
Gemini 3 Pro Preview
Google
77
8786608931---
10
Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview
Google
77
87876589774627-
Page 1 of 7

Cost$20.00
Latency4.02s
Context1M
TTFT4.02s
LegalBench
89
FinanceReasoning
90
Text-to-SQL
90
Agentic RAG
98
Agentic LLM Benchmark
69

Cost$6.00
Latency2.93s
Context1M
TTFT2.93s
LegalBench
86
FinanceReasoning
88
Text-to-SQL
77
Agentic RAG
93
Agentic LLM Benchmark
73

Cost$11.25
Latency4.31s
Context1M
TTFT4.31s
LegalBench
87
FinanceReasoning
90
Text-to-SQL
74
Agentic RAG
87
ARC-AGI-2
93
Agentic LLM Benchmark
62

Cost$3.00
Latency7.25s
Context500k
TTFT7.25s
LegalBench
86
FinanceReasoning
88
Text-to-SQL
79
Agentic RAG
83
Agentic LLM Benchmark
73

Cost$11.25
Latency0.70s
Context272k
TTFT0.70s
LegalBench
87
ARC-AGI-2
85
Agentic LLM Benchmark
59

Cost$4.50
Latency1.62s
Context1M
TTFT1.62s
LegalBench
85
FinanceReasoning
87
Text-to-SQL
71
Agentic RAG
91
ARC-AGI-2
84
Agentic LLM Benchmark
61

Cost$11.25
Latency11.11s
Context1M
TTFT11.11s
FinanceReasoning
91
Text-to-SQL
79
Agentic RAG
96
Agentic LLM Benchmark
54

Cost$10.00
Latency1.71s
Context1M
TTFT1.71s
FinanceReasoning
88
Text-to-SQL
68
Agentic RAG
80
ARC-AGI-2
69
Agentic LLM Benchmark
72

Cost$4.50
Latency-
Context1M
TTFT
LegalBench
87
FinanceReasoning
86
Text-to-SQL
60
Agentic RAG
89
ARC-AGI-2
31

Cost$4.50
Latency22.02s
Context1M
TTFT22.02s
LegalBench
87
FinanceReasoning
87
Text-to-SQL
65
Agentic RAG
89
ARC-AGI-2
77
Agentic LLM Benchmark
46
FrontierMath
27
Page 1 of 7
Frontier Over Time
Intelligence Index by model release date
Cost vs Performance
Blended cost against Intelligence Index
Model × Benchmark
Top models on selected benchmarks
Methodology

How The Index is Built

The Intelligence Index normalizes each benchmark to a 0-100 scale and averages a model's relative standing across the benchmarks it was evaluated on. Each benchmark below feeds that score.

Recent Updates

Recent Updates

Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.

Moonshot AI

Kimi K3

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-5.6 Sol

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-5.6 Terra

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-5.6 Sol Pro

New model added to the AIMultiple Intelligence Index.

AI Model Release TimelineLast 15 models in 41 days

Key AI model releases from leading providers over the past 41 days.

Provider
Google
Google
21 Jul 2026
Gemini 3.5 Flash Lite +1 moreGemini 3.5 Flash LiteGemini 3.6 Flash
13 Aug 2026
Gemini 3.7 Flash
Meta
Meta
16 Jul 2026
Muse Spark 1.1
05 Aug 2026
Muse Spark 1.2
09 Aug 2026
Muse Glimmer 30B
Anthropic
Anthropic
24 Jul 2026
Claude Opus 5
Moonshot AI
Moonshot AI
16 Jul 2026
Kimi K3
OpenAI
OpenAI
09 Jul 2026
GPT-5.6 Luna +5 moreGPT-5.6 LunaGPT-5.6 Luna ProGPT-5.6 SolGPT-5.6 Sol ProGPT-5.6 TerraGPT-5.6 Terra Pro
X
X
08 Jul 2026
Grok 4.5

Explore LLM Use Cases, Analyses & Benchmarks

Top LLMOps Tools & Compare them to MLOPs

LLM
Feature Comparison
Aug 12

LLMOps platforms handle the operational side of running large language models: deployment, monitoring, evaluation, and cost management. We examined top LLMOps tools, their core features, pricing models, and how they differ from each other to help identify the best fit for various use cases. A breakdown of each metric is provided below: LLMOps platforms support…

Read More
LLMAug 11

LLM Pricing: Top 15+ Providers Compared

LLM pricing spans three orders of magnitude: the cheapest commodity models cost under $0.20 per million tokens, while frontier reasoning tiers launched as high as $262.50. The chart below tracks how launch prices moved: each model sits at its launch date with its launch list price per million tokens, blended at a 3:1 input-to-output ratio,…

LLM
Benchmark
Aug 11

Agentic IT: Can LLMs Design a Benchmark

We gave 12 large language models the job a benchmark team does: invent a benchmark, build it, run four models through it, and report the results. Each did it twice. None of the 24 attempts passed every criterion, and six of the rubric’s checks were passed by none of them. The two topics are text-to-SQL,…

LLM
Benchmark
Aug 9

Text-to-SQL: Comparison of LLM Accuracy

We ran 36 large language models over 759 questions from BIRD-SQL, each model writing SQL against a database it had to identify for itself out of 11 candidates. Every parseable query was executed against the real database and its result set compared with the result set of BIRD’s gold query. Missing, malformed and execution-failing queries…

LLM
Open World Evaluation
Aug 6

LLM Orchestration in 2026: 22 Frameworks and Gateways

Optimizing LLM orchestration is key to improving performance while keeping resource use under control. To evaluate how different orchestration approaches perform in practice, we benchmarked: Discover selected LLM orchestration tools, including developer frameworks and enterprise gateways: LLM Orchestration involves managing and integrating multiple Large Language Models (LLMs) to perform complex tasks efficiently. It ensures smooth…

LLM
Benchmark
Aug 4

HALC-Bench: LLM Hallucination on Long-Context Retrieval Benchmark

HALC-Bench (LLM Hallucination on Long-Context Retrieval Benchmark) measures a large language model’s resistance to fabricating evidence for a metric that does not exist in the target document by using 3 haystacks placed at the beginning, middle, and end of the model’s context window, with 204 questions. claude-fable-5 answered all 204 traps correctly at every haystack…

LLM
Benchmark
Aug 4

Compare Multimodal AI Models on Visual Reasoning

We benchmarked 15 leading multimodal AI models on visual reasoning using 200 visual-based questions. The evaluation consisted of two tracks: 100 chart understanding questions testing data visualization interpretation, and 100 visual logic questions assessing pattern recognition and spatial reasoning. Each question was run 5 times to ensure consistent and reliable results. See our benchmark methodology…

LLM
Benchmark
Aug 4

Audience Simulation: Can LLMs Predict Human Behavior?

In marketing, evaluating how accurately LLMs predict human behavior is crucial for assessing their effectiveness in anticipating audience needs and recognizing the risks of misalignment, ineffective communication, or unintended influence. Audience simulation with LLMs enables the modeling of virtual audiences, helping organizations anticipate reactions to content or products without relying on costly surveys or focus…

LLM
Benchmark
Aug 2

AI Gateways for OpenAI: OpenRouter Alternatives

We benchmarked OpenRouter, SambaNova, TogetherAI, Groq, and AI/ML API across three indicators (first-token latency, total latency, and output-token count), with 300 tests using short prompts (approx. 18 tokens) and long prompts (approx. 203 tokens) for total latency. If you plan to use one of these AI gateways, you can: In this benchmark, we compared OpenRouter,…

LLM
Insight
Jul 16

LLM Fine-Tuning Guide for Enterprises

Follow the links for the specific solutions to your LLM output challenges. If your LLM: The widespread adoption of large language models (LLMs) has improved our ability to process human language. However, their generic training often results in suboptimal performance for specific tasks. To overcome this limitation, fine-tuning methods are employed to tailor LLMs to…

LLM
Insight
Jul 12

LLM VRAM Calculator for Self-Hosting

Self-hosting an LLM means running inference on hardware the operator controls rather than via a third-party API, which changes the cost, data control, and privacy profile. Whether a model runs at all depends on memory. The calculator estimates the VRAM or unified memory a model needs to run locally, based on the model, its precision,…