Services
Contact Us

LLM Benchmarks

One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations. See the methodology.

TOP MODEL
Claude Fable 5
Index 92
Best value
MiniMax M3
$0.42 / 1M
Fastest
GPT OSS 120B
0.23s TTFT
Coverage
64 models
8 benchmarks · 336 eval runs
AIMultiple Intelligence Index

Leaderboard

The highest-scoring models across all benchmarks.

Filter & Sort
#
Model
Index
LegalBench
FinanceReasoning
Text-to-SQL
Agentic RAG
ARC-AGI-2
Agentic LLM Benchmark
FrontierMath
Swe-Bench
1
Claude Fable 5
Claude Fable 5
Anthropic
92
89909098-69--
2
Kimi K3
Kimi K3
Moonshot AI
88
86887793-73--
3
GPT-5.6 Sol
GPT-5.6 Sol
OpenAI
87
879074879362--
4
Grok 4.5
Grok 4.5
X
86
86887983-73--
5
GPT-5.5
GPT-5.5
OpenAI
83
87---8559--
6
GPT-5.6 Terra
GPT-5.6 Terra
OpenAI
80
858771918461--
7
GPT-5.6 Sol Pro
GPT-5.6 Sol Pro
OpenAI
80
-917996-54--
8
Claude Opus 4.6
Claude Opus 4.6
Anthropic
78
-8868806972--
9
Gemini 3 Pro Preview
Gemini 3 Pro Preview
Google
77
8786608931---
10
Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview
Google
77
87876589774627-
Page 1 of 7

Cost$20.00
Latency4.02s
Context1M
TTFT4.02s
LegalBench
89
FinanceReasoning
90
Text-to-SQL
90
Agentic RAG
98
Agentic LLM Benchmark
69

Cost$6.00
Latency2.93s
Context1M
TTFT2.93s
LegalBench
86
FinanceReasoning
88
Text-to-SQL
77
Agentic RAG
93
Agentic LLM Benchmark
73

Cost$11.25
Latency4.31s
Context1M
TTFT4.31s
LegalBench
87
FinanceReasoning
90
Text-to-SQL
74
Agentic RAG
87
ARC-AGI-2
93
Agentic LLM Benchmark
62

Cost$3.00
Latency7.25s
Context500k
TTFT7.25s
LegalBench
86
FinanceReasoning
88
Text-to-SQL
79
Agentic RAG
83
Agentic LLM Benchmark
73

Cost$11.25
Latency0.70s
Context272k
TTFT0.70s
LegalBench
87
ARC-AGI-2
85
Agentic LLM Benchmark
59

Cost$4.50
Latency1.62s
Context1M
TTFT1.62s
LegalBench
85
FinanceReasoning
87
Text-to-SQL
71
Agentic RAG
91
ARC-AGI-2
84
Agentic LLM Benchmark
61

Cost$11.25
Latency11.11s
Context1M
TTFT11.11s
FinanceReasoning
91
Text-to-SQL
79
Agentic RAG
96
Agentic LLM Benchmark
54

Cost$10.00
Latency1.71s
Context1M
TTFT1.71s
FinanceReasoning
88
Text-to-SQL
68
Agentic RAG
80
ARC-AGI-2
69
Agentic LLM Benchmark
72

Cost$4.50
Latency-
Context1M
TTFT
LegalBench
87
FinanceReasoning
86
Text-to-SQL
60
Agentic RAG
89
ARC-AGI-2
31

Cost$4.50
Latency22.02s
Context1M
TTFT22.02s
LegalBench
87
FinanceReasoning
87
Text-to-SQL
65
Agentic RAG
89
ARC-AGI-2
77
Agentic LLM Benchmark
46
FrontierMath
27
Page 1 of 7
Frontier Over Time
Intelligence Index by model release date
Cost vs Performance
Blended cost against Intelligence Index
Model × Benchmark
Top models on selected benchmarks
Methodology

How The Index is Built

The Intelligence Index normalizes each benchmark to a 0-100 scale and averages a model's relative standing across the benchmarks it was evaluated on. Each benchmark below feeds that score.

Recent Updates

Recent Updates

Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.

Moonshot AI

Kimi K3

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-5.6 Sol

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-5.6 Terra

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-5.6 Sol Pro

New model added to the AIMultiple Intelligence Index.

Explore LLM Use Cases, Analyses & Benchmarks

Intelligence Density of 71 LLMs for Smarter & Denser Models

LLM
Benchmark
Aug 16

We tracked 71 LLMs released between February 2023 and May 2026 and collected 10 public benchmarks to measure intelligence density. We divided the capability score by the resource the model consumes (active parameters, training compute, and inference price). To calculate intelligence density, we executed the following steps: See methodology for the scoring approach, and per-resource…

Read More
LLM
Insight
Aug 14

50+ ChatGPT Use Cases with Real Life Examples

ChatGPT reached approximately 1 billion weekly active users in early 2026 roughly 10% of the world’s population.1 OpenAI surpassed $20 billion in annual revenue for 2025, confirmed by CFO Sarah Friar.2 The Anthropic Economic Index distinguishes two modes of use: augmentation, in which a human interacts with AI, and automation, in which AI completes tasks…

LLM
Benchmark
Aug 14

Benchmark of 40+ LLMs in Finance: Claude Fable 5 & GPT-5.6 Sol

We evaluated LLMs on 238 hard questions from the FinanceReasoning benchmark (Tang et al.).18 This subset targets the most challenging financial-reasoning tasks, assessing complex, multi-step quantitative reasoning involving financial concepts and formulas. Our evaluation employed a custom prompt design and scoring criteria of accuracy and token consumption. For a detailed explanation of how these metrics…

LLM
Benchmark
Aug 14

AIM Enterprise: Agentic Enterprise Benchmark

Enterprises use LLMs everyday for their regular tasks. To find most cost efficient LLMs, we designed AIM Enterprise, an agentic enterprise benchmark, where we used 69 enterprise tasks in different categories. Each model delivered one file per task. Scores are relative: the judges rank every answer against the other answers to the same task, so…

LLM
Insight
Aug 13

LLM Scaling Laws: Analysis from AI Researchers

Large language models predict the next token based on patterns learned from text data. The term LLM scaling laws refers to empirical regularities that link model performance to the amount of compute, training data, and model parameters used during training. To understand how these relationships influence modern model design in practice, we reviewed findings from…

LLM
Insight
Aug 13

LLM Observability Tools: Weights & Biases, Langsmith

LLM applications have expanded from single-turn chats into multi-step agents that use tools, query databases, and coordinate with other models, making their behavior harder to interpret. LLM observability provides continuous visibility into these complex workflows, helping organizations monitor quality, detect failures, troubleshoot issues, and manage performance and costs. W&B Weave is Weights & Biases‘ LLM…

LLM
Insight
Aug 13

Large Multimodal Models (LMMs) vs LLMs

We evaluated the performance of Large Multimodal Models (LMMs) in financial reasoning tasks using a carefully selected dataset. By analyzing a subset of high-quality financial samples, we assess the models’ capabilities in processing and reasoning with multimodal data in the financial domain. The methodology section provides detailed insights into the dataset and evaluation framework employed.…

LLM
Insight
Aug 12

ChatGPT for Customer Service: Top 10 Use Cases

ChatGPT has moved from novelty to infrastructure in customer service. Companies are using it to cut response times, handle volume their teams can’t absorb, and reduce the cost of routine interactions. But results vary sharply depending on how it’s implemented. OpenAI launched GPT-5.6, a materially more capable model that is better at instruction-following, reasoning across…

LLM
Benchmark
Aug 12

LLM Latency Benchmark by Use Cases in 2026

We benchmarked 11 top large language models with a total of 1,320 requests, splitting reasoning and non-reasoning models, and measured first-token latency, per-token latency, and overall response time. You can find details on how we measured latency here. We report reasoning and non-reasoning models separately. Reasoning models spend several seconds thinking before the first visible…

LLM
Feature Comparison
Aug 12

Top LLMOps Tools & Compare them to MLOPs

LLMOps platforms handle the operational side of running large language models: deployment, monitoring, evaluation, and cost management. We examined top LLMOps tools, their core features, pricing models, and how they differ from each other to help identify the best fit for various use cases. A breakdown of each metric is provided below: LLMOps platforms support…

LLMAug 11

LLM Pricing: Top 15+ Providers Compared

LLM pricing spans three orders of magnitude: the cheapest commodity models cost under $0.20 per million tokens, while frontier reasoning tiers launched as high as $262.50. The chart below tracks how launch prices moved: each model sits at its launch date with its launch list price per million tokens, blended at a 3:1 input-to-output ratio,…