Premium
Services
Premium

LLM Benchmarks

One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations. How the Index Is Built

TOP MODEL
Claude Opus 5.5
Index 79
Best value
Qwen3.8 Flash
$0.23 / 1M
Fastest
Command A+
0.22s TTFT
AIMultiple Intelligence Index

Leaderboard

The highest-scoring models across all benchmarks.

Filter & Sort
#
Model
Index
AIM-Agentic-RAG
AIM-AI-Bias
AIM-FinanceReasoning
AIM-Text-to-SQL
ARC-AGI-3
CritPt
FrontierMath
IFBench
MMMU-Pro
SciCode
Terminal-Bench 2.1
1
Claude Opus 5.5
Claude Opus 5.5
Anthropic
79
90949288-32--8867-
2
GPT-6 Astra
GPT-6 Astra
OpenAI
76
84988766633298-875690
3
Claude Fable 5.1
Claude Fable 5.1
Anthropic
73
84-9067-3188--6391
4
Qwen3.8 Flash
Qwen3.8 Flash
Alibaba Cloud
71
-------81-4786
5
Qwen3.8 Max
Qwen3.8 Max
Alibaba Cloud
70
8393-51--46-825381
6
Claude Opus 5
Claude Opus 5
Anthropic
68
85819064302973-855689
7
Muse Spark 1.1
Muse Spark 1.1
Meta
68
---------5878
8
GPT-5.6 Sol
GPT-5.6 Sol
OpenAI
67
778990548328373835790
9
Kimi K3
Kimi K3
Moonshot AI
66
83898848-2339-815985
10
Claude Sonnet 5
Claude Sonnet 5
Anthropic
65
73918738--29--5480
Page 1 of 4

Cost$8.00
Latency1.16s
Context1M
TTFT1.16s
AIM-Agentic-RAG
90
AIM-AI-Bias
94
AIM-FinanceReasoning
92
AIM-Text-to-SQL
88
CritPt
32
MMMU-Pro
88
SciCode
67

Cost$20.00
Latency4.75s
Context524k
TTFT4.75s
AIM-Agentic-RAG
84
AIM-AI-Bias
98
AIM-FinanceReasoning
87
AIM-Text-to-SQL
66
ARC-AGI-3
63
CritPt
32
FrontierMath
98
MMMU-Pro
87
SciCode
56
Terminal-Bench 2.1
90

Cost$20.00
Latency1.93s
Context1M
TTFT1.93s
AIM-Agentic-RAG
84
AIM-FinanceReasoning
90
AIM-Text-to-SQL
67
CritPt
31
FrontierMath
88
SciCode
63
Terminal-Bench 2.1
91

Cost$0.23
Latency1.72s
Context1M
TTFT1.72s
IFBench
81
SciCode
47
Terminal-Bench 2.1
86

Cost$3.00
Latency56.57s
Context1M
TTFT56.57s
AIM-Agentic-RAG
83
AIM-AI-Bias
93
AIM-Text-to-SQL
51
FrontierMath
46
MMMU-Pro
82
SciCode
53
Terminal-Bench 2.1
81

Cost$10.00
Latency2.58s
Context1M
TTFT2.58s
AIM-Agentic-RAG
85
AIM-AI-Bias
81
AIM-FinanceReasoning
90
AIM-Text-to-SQL
64
ARC-AGI-3
30
CritPt
29
FrontierMath
73
MMMU-Pro
85
SciCode
56
Terminal-Bench 2.1
89

Cost$2.00
Latency1.47s
Context1M
TTFT1.47s
SciCode
58
Terminal-Bench 2.1
78

Cost$8.00
Latency3.06s
Context1M
TTFT3.06s
AIM-Agentic-RAG
77
AIM-AI-Bias
89
AIM-FinanceReasoning
90
AIM-Text-to-SQL
54
ARC-AGI-3
8
CritPt
32
FrontierMath
83
IFBench
73
MMMU-Pro
83
SciCode
57
Terminal-Bench 2.1
90

Cost$3.00
Latency2.70s
Context1M
TTFT2.70s
AIM-Agentic-RAG
83
AIM-AI-Bias
89
AIM-FinanceReasoning
88
AIM-Text-to-SQL
48
CritPt
23
FrontierMath
39
MMMU-Pro
81
SciCode
59
Terminal-Bench 2.1
85

Cost$4.00
Latency1.79s
Context1M
TTFT1.79s
AIM-Agentic-RAG
73
AIM-AI-Bias
91
AIM-FinanceReasoning
87
AIM-Text-to-SQL
38
FrontierMath
29
SciCode
54
Terminal-Bench 2.1
80
Page 1 of 4
Frontier Over Time
Intelligence Index by model release date
Cost vs Performance
Blended cost against Intelligence Index
Model × Benchmark
Top models on selected benchmarks

How the Index Is Built

The index is the average of a model's scores on the benchmarks listed below, over the ones it has actually run: a benchmark it was never run on is not counted as a zero, it is simply absent. Every benchmark in the index counts equally. A benchmark that already reports a 0-100 percentage is used as it stands; one on a different scale is rescaled onto 0-100 first. A model needs results on at least two index benchmarks before it is ranked, so a single result does not give a model a standing of its own. Because scores are used at face value, a benchmark the whole field finds hard pulls every score on it down, and a model measured on harder benchmarks carries that into its average.

Benchmarks We Used

Recent Updates

Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.

Anthropic

Claude Opus 5.5

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-6 Sol

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-6 Luna

New model added to the AIMultiple Intelligence Index.

X

Grok 4.7

New model added to the AIMultiple Intelligence Index.

LLM Release TimelineThe pace of AI innovation
AI Labs
Alibaba Cloud
Alibaba Cloud
27 Jul 2026
Qwen3.7 Flash
03 Aug 2026
Qwen3.8 Max
12 Aug 2026
Qwen3.8 2.4T A95B
14 Aug 2026
Qwen3.8 27B
26 Aug 2026
Qwen3.8 Flash
03 Sep 2026
Qwen3.8 Max (0902)
21 Sep 2026
Qwen3.8 Omni Flash
23 Sep 2026
Qwen3.8 Max Prime
Z AI
Z AI
18 Aug 2026
GLM 5.3
26 Aug 2026
GLM 5.3 Flash
18 Sep 2026
GLM 5.3 FlashX
23 Sep 2026
GLM 5.3 Prime
Anthropic
Anthropic
30 Jun 2026
Claude Sonnet 5
24 Jul 2026
Claude Opus 5
01 Sep 2026
Claude Fable 5.1
22 Sep 2026
Claude Opus 5.5
OpenAI
OpenAI
09 Jul 2026
GPT-5.6 Luna +2 moreGPT-5.6 LunaGPT-5.6 SolGPT-5.6 Terra
04 Sep 2026
GPT-6 Astra
22 Sep 2026
GPT-6 Luna +1 moreGPT-6 LunaGPT-6 Sol
Cohere
Cohere
22 Sep 2026
Command A+
X
X
08 Jul 2026
Grok 4.5
12 Aug 2026
Grok 4.6
21 Sep 2026
Grok 4.7
DeepSeek
DeepSeek
31 Jul 2026
DeepSeek V4 Flash 0731
12 Aug 2026
DeepSeek V4 Pro 0813
21 Aug 2026
DeepSeek V4 Flash Vision Exp
10 Sep 2026
DeepSeek V4.1 Flash
Meta
Meta
16 Jul 2026
Muse Spark 1.1
05 Aug 2026
Muse Spark 1.2
09 Aug 2026
Muse Glimmer 30B
21 Aug 2026
Muse Spark 1.2 Contributor
02 Sep 2026
Muse Spark 1.3 +1 moreMuse Spark 1.3Muse Spark 1.3 Contributor
Google
Google
30 Jun 2026
Nano Banana 2 Lite
21 Jul 2026
Gemini 3.5 Flash Lite +1 moreGemini 3.5 Flash LiteGemini 3.6 Flash
13 Aug 2026
Gemini 3.7 Flash
02 Sep 2026
Gemini 3.8 Flash
Tencent
Tencent
06 Jul 2026
Hy3
19 Aug 2026
Hy-MT2-7B
20 Aug 2026
Hy-MT2-1.8B +1 moreHy-MT2-1.8BHy-MT2-30B-A3B
28 Aug 2026
Hy4 +1 moreHy4Hy4 preview
Moonshot AI
Moonshot AI
16 Jul 2026
Kimi K3

Explore LLM Use Cases, Analyses & Benchmarks

Compare 9 Large Language Models in Healthcare

LLM
Feature Comparison
Sep 15

We benchmarked 9 LLMs using the MedQA dataset, a graduate-level clinical exam benchmark derived from USMLE questions. Each model answered the same multiple-choice clinical scenarios using a standardized prompt, enabling direct comparison of accuracy. We also recorded latency per question by dividing total runtime by the number of MedQA items completed. Benchmark methodology: This benchmark…

Read More
LLM
Benchmark
Sep 15

Audience Simulation: Can LLMs Predict Human Behavior?

In marketing, evaluating how accurately LLMs predict human behavior is crucial for assessing their effectiveness in anticipating audience needs and recognizing the risks of misalignment, ineffective communication, or unintended influence. Audience simulation with LLMs enables the modeling of virtual audiences, helping organizations anticipate reactions to content or products without relying on costly surveys or focus…

LLM
Benchmark
Sep 15

HALC-Bench: LLM Hallucination on Long-Context Retrieval Benchmark

HALC-Bench (LLM Hallucination on Long-Context Retrieval Benchmark) measures a large language model’s resistance to fabricating evidence for a metric that does not exist in the target document by using 3 haystacks placed at the beginning, middle, and end of the model’s context window, with 204 questions. claude-fable-5 answered all 204 traps correctly at every haystack…

LLM
Benchmark
Sep 15

AI Gateways for OpenAI: OpenRouter Alternatives

We benchmarked OpenRouter, SambaNova, TogetherAI, Groq, and AI/ML API across three indicators (first-token latency, total latency, and output-token count), with 300 tests using short prompts (approx. 18 tokens) and long prompts (approx. 203 tokens) for total latency. If you plan to use one of these AI gateways, you can: In this benchmark, we compared OpenRouter,…

LLM
Insight
Sep 2

LLM Observability Tools: Weights & Biases, Langsmith

LLM applications have expanded from single-turn chats into multi-step agents that use tools, query databases, and coordinate with other models, making their behavior harder to interpret. LLM observability provides continuous visibility into these complex workflows, helping organizations monitor quality, detect failures, troubleshoot issues, and manage performance and costs. W&B Weave is Weights & Biases‘ LLM…

LLM
Feature Comparison
Sep 2

Cloud LLM vs Local LLMs: Examples & Benefits

Cloud LLMs, powered by advanced models like GPT-5.5 and Claude Opus 4.7, offer scalability and accessibility. Conversely, Local LLMs, driven by open-source models such as Llama 4, DeepSeek V4, and Qwen3.6-Plus, ensure stronger privacy and customization. Explore what are cloud LLMs, strengths and weaknesses, most common case studies with real-life examples, and how they differ…

LLM
Insight
Sep 1

LLM Fine-Tuning Guide for Enterprises

Follow the links for the specific solutions to your LLM output challenges. If your LLM: The widespread adoption of large language models (LLMs) has improved our ability to process human language. However, their generic training often results in suboptimal performance for specific tasks. To overcome this limitation, fine-tuning methods are employed to tailor LLMs to…

LLM
Insight
Sep 1

10+ Large Language Model Examples

We have gathered open-source benchmarks to compare leading proprietary and open-source large language models. Choose your use case to find the right model. You can evaluate large language models by examining their benchmark performance and real-world latency (available by clicking each model’s name in the table), and by reviewing their pricing to assess overall efficiency…

LLM
Feature Comparison
Aug 27

LLM Pricing: Top 15+ Providers Compared

LLM pricing spans four orders of magnitude: the cheapest models launched under $0.03 per million tokens, while frontier reasoning tiers launched at up to $262.50. The chart below tracks launch prices: each point is the average price of the models one size class launched in a calendar quarter, blended 3 parts input to 1 part…

LLM
Insight
Aug 27

LLM Automation: Top 7 Tools & 8 Case Studies 

LLM automation refers to shift to intelligent automation tools that leverage LLMs, including AI agents, fine-tuned LLMs and RAG models to automate and coordinate tasks. Explore what LLM automation is, its top real-life applications and major tools: Large language models in automation is a systematic approach that combines Natural Language Processing (NLP) with existing process…

LLM
Open World Evaluation
Aug 26

LLM Orchestration: 22 Frameworks and Gateways

Optimizing LLM orchestration is key to improving performance while keeping resource use under control. To evaluate how different orchestration approaches perform in practice, we benchmarked: Discover selected LLM orchestration tools, including developer frameworks and enterprise gateways: LLM Orchestration involves managing and integrating multiple Large Language Models (LLMs) to perform complex tasks efficiently. It ensures smooth…