Premium
Services
Premium

LLM Benchmarks

One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations. How the Index Is Built

TOP MODEL
Claude Opus 5.5
Index 76
Best value
Qwen3.8 Flash
$0.23 / 1M
Fastest
Command A+
0.22s TTFT
AIMultiple Intelligence Index

Leaderboard

The highest-scoring models across all benchmarks.

Filter & Sort
#
Model
Index
AIM-Agentic-RAG
AIM-FinanceReasoning
AIM-Text-to-SQL
ARC-AGI-3
CritPt
FrontierMath
IFBench
MMMU-Pro
SciCode
Terminal-Bench 2.1
1
Claude Opus 5.5
Claude Opus 5.5
Anthropic
76
909288-32--8867-
2
GPT-6 Astra
GPT-6 Astra
OpenAI
74
848766633298-875690
3
Claude Fable 5.1
Claude Fable 5.1
Anthropic
73
849067-3188--6391
4
Qwen3.8 Flash
Qwen3.8 Flash
Alibaba Cloud
71
------81-4786
5
Muse Spark 1.1
Muse Spark 1.1
Meta
68
--------5878
6
Claude Opus 5
Claude Opus 5
Anthropic
67
859064302973-855689
7
Qwen3.8 Max
Qwen3.8 Max
Alibaba Cloud
66
83-51--46-825381
8
GPT-5.6 Sol
GPT-5.6 Sol
OpenAI
65
7790548328373835790
9
Kimi K3
Kimi K3
Moonshot AI
63
838848-2339-815985
10
Gemini 3.7 Flash
Gemini 3.7 Flash
Google
63
85-72-1437-866086
Page 1 of 4

Cost$8.00
Latency1.16s
Context1M
TTFT1.16s
AIM-Agentic-RAG
90
AIM-FinanceReasoning
92
AIM-Text-to-SQL
88
CritPt
32
MMMU-Pro
88
SciCode
67

Cost$20.00
Latency4.75s
Context524k
TTFT4.75s
AIM-Agentic-RAG
84
AIM-FinanceReasoning
87
AIM-Text-to-SQL
66
ARC-AGI-3
63
CritPt
32
FrontierMath
98
MMMU-Pro
87
SciCode
56
Terminal-Bench 2.1
90

Cost$20.00
Latency1.93s
Context1M
TTFT1.93s
AIM-Agentic-RAG
84
AIM-FinanceReasoning
90
AIM-Text-to-SQL
67
CritPt
31
FrontierMath
88
SciCode
63
Terminal-Bench 2.1
91

Cost$0.23
Latency1.72s
Context1M
TTFT1.72s
IFBench
81
SciCode
47
Terminal-Bench 2.1
86

Cost$2.00
Latency1.47s
Context1M
TTFT1.47s
SciCode
58
Terminal-Bench 2.1
78

Cost$10.00
Latency2.58s
Context1M
TTFT2.58s
AIM-Agentic-RAG
85
AIM-FinanceReasoning
90
AIM-Text-to-SQL
64
ARC-AGI-3
30
CritPt
29
FrontierMath
73
MMMU-Pro
85
SciCode
56
Terminal-Bench 2.1
89

Cost$3.00
Latency56.57s
Context1M
TTFT56.57s
AIM-Agentic-RAG
83
AIM-Text-to-SQL
51
FrontierMath
46
MMMU-Pro
82
SciCode
53
Terminal-Bench 2.1
81

Cost$8.00
Latency3.06s
Context1M
TTFT3.06s
AIM-Agentic-RAG
77
AIM-FinanceReasoning
90
AIM-Text-to-SQL
54
ARC-AGI-3
8
CritPt
32
FrontierMath
83
IFBench
73
MMMU-Pro
83
SciCode
57
Terminal-Bench 2.1
90

Cost$3.00
Latency2.70s
Context1M
TTFT2.70s
AIM-Agentic-RAG
83
AIM-FinanceReasoning
88
AIM-Text-to-SQL
48
CritPt
23
FrontierMath
39
MMMU-Pro
81
SciCode
59
Terminal-Bench 2.1
85

Cost$0.76
Latency6.15s
Context1M
TTFT6.15s
AIM-Agentic-RAG
85
AIM-Text-to-SQL
72
CritPt
14
FrontierMath
37
MMMU-Pro
86
SciCode
60
Terminal-Bench 2.1
86
Page 1 of 4
Frontier Over Time
Intelligence Index by model release date
Cost vs Performance
Blended cost against Intelligence Index
Model × Benchmark
Top models on selected benchmarks

How the Index Is Built

The index is the average of a model's scores on the benchmarks listed below, over the ones it has actually run: a benchmark it was never run on is not counted as a zero, it is simply absent. Every benchmark in the index counts equally. A benchmark that already reports a 0-100 percentage is used as it stands; one on a different scale is rescaled onto 0-100 first. A model needs results on at least two index benchmarks before it is ranked, so a single result does not give a model a standing of its own. Because scores are used at face value, a benchmark the whole field finds hard pulls every score on it down, and a model measured on harder benchmarks carries that into its average.

Benchmarks We Used

Recent Updates

Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.

Anthropic

Claude Opus 5.5

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-6 Sol

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-6 Luna

New model added to the AIMultiple Intelligence Index.

X

Grok 4.7

New model added to the AIMultiple Intelligence Index.

LLM Release TimelineThe pace of AI innovation
AI Labs
Alibaba Cloud
Alibaba Cloud
27 Jul 2026
Qwen3.7 Flash
03 Aug 2026
Qwen3.8 Max
12 Aug 2026
Qwen3.8 2.4T A95B
14 Aug 2026
Qwen3.8 27B
26 Aug 2026
Qwen3.8 Flash
03 Sep 2026
Qwen3.8 Max (0902)
21 Sep 2026
Qwen3.8 Omni Flash
23 Sep 2026
Qwen3.8 Max Prime
Z AI
Z AI
18 Aug 2026
GLM 5.3
26 Aug 2026
GLM 5.3 Flash
18 Sep 2026
GLM 5.3 FlashX
23 Sep 2026
GLM 5.3 Prime
Anthropic
Anthropic
30 Jun 2026
Claude Sonnet 5
24 Jul 2026
Claude Opus 5
01 Sep 2026
Claude Fable 5.1
22 Sep 2026
Claude Opus 5.5
OpenAI
OpenAI
09 Jul 2026
GPT-5.6 Luna +2 moreGPT-5.6 LunaGPT-5.6 SolGPT-5.6 Terra
04 Sep 2026
GPT-6 Astra
22 Sep 2026
GPT-6 Luna +1 moreGPT-6 LunaGPT-6 Sol
Cohere
Cohere
22 Sep 2026
Command A+
X
X
08 Jul 2026
Grok 4.5
12 Aug 2026
Grok 4.6
21 Sep 2026
Grok 4.7
DeepSeek
DeepSeek
31 Jul 2026
DeepSeek V4 Flash 0731
12 Aug 2026
DeepSeek V4 Pro 0813
21 Aug 2026
DeepSeek V4 Flash Vision Exp
10 Sep 2026
DeepSeek V4.1 Flash
Meta
Meta
16 Jul 2026
Muse Spark 1.1
05 Aug 2026
Muse Spark 1.2
09 Aug 2026
Muse Glimmer 30B
21 Aug 2026
Muse Spark 1.2 Contributor
02 Sep 2026
Muse Spark 1.3 +1 moreMuse Spark 1.3Muse Spark 1.3 Contributor
Google
Google
30 Jun 2026
Nano Banana 2 Lite
21 Jul 2026
Gemini 3.5 Flash Lite +1 moreGemini 3.5 Flash LiteGemini 3.6 Flash
13 Aug 2026
Gemini 3.7 Flash
02 Sep 2026
Gemini 3.8 Flash
Tencent
Tencent
06 Jul 2026
Hy3
19 Aug 2026
Hy-MT2-7B
20 Aug 2026
Hy-MT2-1.8B +1 moreHy-MT2-1.8BHy-MT2-30B-A3B
28 Aug 2026
Hy4 +1 moreHy4Hy4 preview
Moonshot AI
Moonshot AI
16 Jul 2026
Kimi K3

Explore LLM Use Cases, Analyses & Benchmarks

Text-to-SQL Benchmark: SQL Accuracy Across 40+ LLMs

LLM
Benchmark
Sep 25

SQL accuracy is the percentage of scored queries that return the reference result. Incorrect routes and references flagged as broken are excluded. A reference query is the SQL supplied as the expected answer. Each model reaches a different set of SQL questions because scoring depends on its database choices. Differences in these subsets and execution…

Read More
LLM
Insight
Sep 25

LLM Market Share: Compare Usage & Adoption

We analyzed LLM market share by combining usage-based data and web visit estimates to show how demand for large language models is distributed across AI labs and AI applications. Read the methodology to see how we measured and calculated these results. The United States dominated web visits across all four months, consistently accounting for 85–95%.…

LLM
Benchmark
Sep 24

Benchmark of 40+ LLMs in Finance: Claude Opus 5.5 & GPT-6 Astra

We evaluated LLMs on 238 hard questions from the FinanceReasoning benchmark (Tang et al.).6 This subset targets the most challenging financial-reasoning tasks, assessing complex, multi-step quantitative reasoning involving financial concepts and formulas. Our evaluation employed a custom prompt design and scoring criteria of accuracy and token consumption. For a detailed explanation of how these metrics…

LLM
Benchmark
Sep 21

Compare Multimodal AI Models on Visual Reasoning

We benchmarked 15 leading multimodal AI models on visual reasoning using 200 visual-based questions. The evaluation consisted of two tracks: 100 chart understanding questions testing data visualization interpretation, and 100 visual logic questions assessing pattern recognition and spatial reasoning. Each question was run 5 times to ensure consistent and reliable results. See our benchmark methodology…

LLM
Insight
Sep 21

Large Multimodal Models (LMMs) vs LLMs

Evaluate LLMs and LMMs by comparing their benchmark scores and real-world latency by clicking the model’s name in the table below. You can also weigh their input and output pricing to judge overall efficiency and value. *Audio is native on the E2B, E4B and 12B models only. Available in five sizes: E2B, E4B, 12B, 26B…

LLM
Insight
Sep 21

LLM VRAM Calculator for Self-Hosting

Self-hosting an LLM means running inference on hardware the operator controls rather than via a third-party API, which changes the cost, data control, and privacy profile. Whether a model runs at all depends on memory. The calculator estimates the VRAM or unified memory a model needs to run locally, based on the model, its precision,…

LLM
Benchmark
Sep 18

Agentic IT: Can AI Agents Design a Benchmark

We tested 16 models on a benchmark design in text-to-SQL and tool calling. Each model built one benchmark per topic, for a total of 32 submissions. No submission passed every rubric criterion. The agents could build and run tests, but none demonstrated both a blank-answer test and a correct-answer test of their own scorer. Text-to-SQL…

LLM
Benchmark
Sep 17

AIM Enterprise: Agentic Enterprise Benchmark

Enterprises use LLMs every day for their regular tasks. To find the most cost-efficient LLMs, we designed AIM Enterprise, an agentic enterprise benchmark, where we used 69 real enterprise tasks across strategy, marketing, HR, sales, and operations. Two judge models scored every file that passed the checks. On 32.7% of the individual scores the two…

LLM
Insight
Sep 17

The Future of Large Language Models

See the future of large language models by delving into promising approaches, such as self-training, fact-checking, and sparse expertise that could address LLM limitations. Success rate comparison of LLM’s Claude Sonnet 4.6 led the benchmark with an overall score of 0.748, with base and thinking variants tied to three decimal places. Claude Opus 4.8 (0.702),…

LLM
Benchmark
Sep 15

Large Language Models in Cybersecurity

We evaluated 7 large language models across 9 cybersecurity domains using SecBench, a large-scale and multi-format benchmark for security tasks. We tested each model on 44,823 multiple-choice questions (MCQs) and 3,087 short-answer questions (SAQs), covering data security, identity & access management, network security, vulnerability management, and cloud security. MCQs (Multiple-Choice Questions) benchmarking: SAQs (Short Answer…

LLM
Insight
Sep 15

LLM Scaling Laws: Analysis from AI Researchers

Large language models predict the next token based on patterns learned from text data. The term LLM scaling laws refers to empirical regularities that link model performance to the amount of compute, training data, and model parameters used during training. To understand how these relationships influence modern model design in practice, we reviewed findings from…