Services
Contact Us

LLM Benchmarks

One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations.

Last Updated Jun 2026
TOP MODEL
Claude Fable 5
Index 92
Best value
MiniMax M3
$0.42 / 1M
Fastest
GPT OSS 120B
0.19s TTFT
Coverage
64 models
8 benchmarks · 336 eval runs
AIMultiple Intelligence Index

Leaderboard

The highest-scoring models across all benchmarks.

Filter & Sort
#
Model
Index
LegalBench
FinanceReasoning
Text-to-SQL
Agentic RAG
ARC-AGI-2
Agentic LLM Benchmark
FrontierMath
Swe-Bench
1
Claude Fable 5
Claude Fable 5
Anthropic
92
89909098-69--
2
Kimi K3
Kimi K3
Moonshot AI
88
86887793-73--
3
GPT-5.6 Sol
GPT-5.6 Sol
OpenAI
87
879074879362--
4
Grok 4.5
Grok 4.5
X
86
86887983-73--
5
GPT-5.5
GPT-5.5
OpenAI
83
87---8559--
6
GPT-5.6 Terra
GPT-5.6 Terra
OpenAI
80
858771918461--
7
GPT-5.6 Sol Pro
GPT-5.6 Sol Pro
OpenAI
80
-917996-54--
8
Claude Opus 4.6
Claude Opus 4.6
Anthropic
78
-8868806972--
9
Gemini 3 Pro Preview
Gemini 3 Pro Preview
Google
77
8786608931---
10
Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview
Google
77
87876589774627-
Page 1 of 7

Cost$20.00
Latency3.98s
Context1M
TTFT3.98s
LegalBench
89
FinanceReasoning
90
Text-to-SQL
90
Agentic RAG
98
Agentic LLM Benchmark
69

Cost$6.00
Latency252.11s
Context1M
TTFT252.11s
LegalBench
86
FinanceReasoning
88
Text-to-SQL
77
Agentic RAG
93
Agentic LLM Benchmark
73

Cost$11.25
Latency2.47s
Context1M
TTFT2.47s
LegalBench
87
FinanceReasoning
90
Text-to-SQL
74
Agentic RAG
87
ARC-AGI-2
93
Agentic LLM Benchmark
62

Cost$3.00
Latency10.26s
Context500k
TTFT10.26s
LegalBench
86
FinanceReasoning
88
Text-to-SQL
79
Agentic RAG
83
Agentic LLM Benchmark
73

Cost$11.25
Latency1.03s
Context272k
TTFT1.03s
LegalBench
87
ARC-AGI-2
85
Agentic LLM Benchmark
59

Cost$5.63
Latency1.92s
Context1M
TTFT1.92s
LegalBench
85
FinanceReasoning
87
Text-to-SQL
71
Agentic RAG
91
ARC-AGI-2
84
Agentic LLM Benchmark
61

Cost$11.25
Latency11.46s
Context1M
TTFT11.46s
FinanceReasoning
91
Text-to-SQL
79
Agentic RAG
96
Agentic LLM Benchmark
54

Cost$10.00
Latency1.75s
Context1M
TTFT1.75s
FinanceReasoning
88
Text-to-SQL
68
Agentic RAG
80
ARC-AGI-2
69
Agentic LLM Benchmark
72

Cost$4.50
Latency-
Context1M
TTFT
LegalBench
87
FinanceReasoning
86
Text-to-SQL
60
Agentic RAG
89
ARC-AGI-2
31

Cost$4.50
Latency33.83s
Context1M
TTFT33.83s
LegalBench
87
FinanceReasoning
87
Text-to-SQL
65
Agentic RAG
89
ARC-AGI-2
77
Agentic LLM Benchmark
46
FrontierMath
27
Page 1 of 7

Charts

Following "Best Performing" models

Frontier Over Time
Intelligence Index by model release date
Cost vs Performance
Blended cost against Intelligence Index
Model × Benchmark
Top models on selected benchmarks
Methodology

How The Index is Built

The Intelligence Index normalizes each benchmark to a 0-100 scale and averages a model's relative standing across the benchmarks it was evaluated on. Each benchmark below feeds that score.

Recent Updates

Recent Updates

Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.

Moonshot AI

Kimi K3

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-5.6 Sol

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-5.6 Terra

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-5.6 Sol Pro

New model added to the AIMultiple Intelligence Index.

Explore LLM Use Cases, Analyses & Benchmarks

LLM Scaling Laws: Analysis from AI Researchers

LLM
Insight
Jul 24

Large language models predict the next token based on patterns learned from text data. The term LLM scaling laws refers to empirical regularities that link model performance to the amount of compute, training data, and model parameters used during training. To understand how these relationships influence modern model design in practice, we reviewed findings from…

Read More
LLMJul 23

LLM Pricing: Top 15+ Providers Compared

LLM pricing spans three orders of magnitude: the cheapest commodity models cost under $0.20 per million tokens, while frontier reasoning tiers launched as high as $262.50. The chart below tracks how launch prices moved: each model sits at its launch date with its launch list price per million tokens, blended at a 3:1 input-to-output ratio,…

LLM
Benchmark
Jul 17

Text-to-SQL: Comparison of LLM Accuracy

I have relied on SQL for data analysis for 18 years, beginning in my days as a consultant. Translating natural-language questions into SQL makes data more accessible, allowing anyone, even those without technical skills, to work directly with databases. We used our text-to-SQL benchmark methodology on 35+ large language models (LLMs) to assess their performance…

LLM
Insight
Jul 16

LLM Fine-Tuning Guide for Enterprises

Follow the links for the specific solutions to your LLM output challenges. If your LLM: The widespread adoption of large language models (LLMs) has improved our ability to process human language. However, their generic training often results in suboptimal performance for specific tasks. To overcome this limitation, fine-tuning methods are employed to tailor LLMs to…

LLM
Insight
Jul 16

LLM Observability Tools: Weights & Biases, Langsmith

LLM applications have expanded from single-turn chats into multi-step agents that use tools, query databases, and coordinate with other models, making their behavior harder to interpret. LLM observability provides continuous visibility into these complex workflows, helping organizations monitor quality, detect failures, troubleshoot issues, and manage performance and costs. W&B Weave is Weights & Biases‘ LLM…

LLM
Insight
Jul 12

LLM VRAM Calculator for Self-Hosting

Self-hosting an LLM means running inference on hardware the operator controls rather than via a third-party API, which changes the cost, data control, and privacy profile. Whether a model runs at all depends on memory. The calculator estimates the VRAM or unified memory a model needs to run locally, based on the model, its precision,…

LLM
Benchmark
Jul 10

Benchmark of 40+ LLMs in Finance: Claude Fable 5 & GPT-5.6 Sol

We evaluated 40+ LLMs in finance on 238 hard questions from the FinanceReasoning benchmark to identify which models excel at complex financial reasoning tasks like statement analysis, forecasting, and ratio calculations. We evaluated LLMs on 238 hard questions from the FinanceReasoning benchmark (Tang et al.).81 This subset targets the most challenging financial-reasoning tasks, assessing complex,…

LLM
Insight
Jul 10

LLM Automation: Top 7 Tools & 8 Case Studies 

LLM automation refers to shift to intelligent automation tools that leverage LLMs, including AI agents, fine-tuned LLMs and RAG models to automate and coordinate tasks. Explore what LLM automation is, its top real-life applications and major tools: Large language models in automation is a systematic approach that combines Natural Language Processing (NLP) with existing process…

LLM
Benchmark
Jul 8

LLM Latency Benchmark by Use Cases in 2026

We benchmarked 11 top large language models with a total of 1,320 requests, splitting reasoning and non-reasoning models, and measured first-token latency, per-token latency, and overall response time. You can find details on how we measured latency here. We report reasoning and non-reasoning models separately. Reasoning models spend several seconds thinking before the first visible…

LLM
Benchmark
Jul 7

HALC-Bench: LLM Hallucination on Long-Context Retrieval Benchmark

HALC-Bench (LLM Hallucination on Long-Context Retrieval Benchmark) measures a large language model’s resistance to fabricating evidence for a metric that does not exist in the target document by using 3 haystacks placed at the beginning, middle, and end of the model’s context window, with 204 questions. claude-fable-5 answered all 204 traps correctly at every haystack…

LLM
Benchmark
Jul 7

Intelligence Density of 71 LLMs: Smarter and Denser Models

We tracked 71 LLMs released between February 2023 and May 2026 and collected 10 public benchmarks to measure intelligence density. We divided the capability score by the resource the model consumes (active parameters, training compute, and inference price). To calculate intelligence density, we executed the following steps: See methodology for the scoring approach, and per-resource…