Premium
Services
Premium

LLM Benchmarks

One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations. How the Index Is Built

TOP MODEL
Claude Opus 5.5
Index 100
Best value
Qwen3.8 Flash
$0.23 / 1M
Fastest
MiniMax M2.7
0.15s TTFT
AIMultiple Intelligence Index

Leaderboard

The highest-scoring models across all benchmarks.

Filter & Sort
#
Model
Index
AA-LCR
AIM-A-CODE-LLM Bench
AIM-Agentic-RAG
AIM-FinanceReasoning
AIM-HALC-Bench
AIM-RELC-Bench
AIM-Text-to-SQL
APEX-Agents-AA
ARC-AGI-1
ARC-AGI-2
ARC-AGI-3
ArxivMath
AutomationBench-AA
BioMysteryBench
BrowseComp
Convex Coding Evals
CritPt
Crosby RedlineBench
DeepSWE 1.1
EnterpriseOps-Gym-AA
Frontier-Bench
FrontierCode
FrontierMath
GDPval
GPQA Diamond
Harvey LAB-AA
HealthBench Professional
HieroglyphBench
Humanity's Last Exam
IFBench
ITBench-AA
LegalBench
LiveCodeBench
MMMU-Pro
ObviousBench
Opus Magnum Bench
ReactBench
RuneBench
SciCode
SimpleBench
Swe-Bench
Tau2-Bench Telecom
Tau3-Banking
Terminal-Bench 2.1
Terminal-Bench 4.0
Terminal-Bench Hard
Terminal-Bench-Science
1
Claude Opus 5.5
Claude Opus 5.5
Anthropic
100
--9092--88---------32------1846--------------67--------
2
GPT-6 Astra
GPT-6 Astra
OpenAI
92
--8487--66--9563-41-928532-74--5398154296-70----------756----9058--
3
Claude Fable 5.1
Claude Fable 5.1
Anthropic
91
80-8490--67-9890--31--8131-67--5188173594-62-59--89-----663---479156--
4
Qwen3.8 Flash
Qwen3.8 Flash
Alibaba Cloud
87
77----------------------174392---3881--------47---4586---
5
Command A+
Command A+
Cohere
87
----------------30------------74-61---------81---25-
6
MiMo-V2.6-Pro
MiMo-V2.6-Pro
Xiaomi
87
----------------27------1673--------------61--------
7
Muse Spark 1.1
Muse Spark 1.1
Meta
86
81-----------------53----138190-59-46--85--99-23658---3278---
8
Claude Fable 5
Claude Fable 5
Anthropic
85
77698290--66599989-7917-878429-70-3454901741931463235664-89-8199224766182-993885426321
9
Claude Opus 5
Claude Opus 5
Anthropic
83
79-8590--64-989030915079918229-74-445373170894-60-55--87-85100--656---458952-30
10
GPT-6 Sol
GPT-6 Sol
OpenAI
79
--82---54---------31------1487--------------58--------
Page 1 of 13

Cost$8.00
Latency2.81s
Context1M
TTFT2.81s
AIM-Agentic-RAG
90
AIM-FinanceReasoning
92
AIM-Text-to-SQL
88
CritPt
32
GDPval
1846
SciCode
67

Cost$20.00
Latency6.31s
Context1M
TTFT6.31s
AIM-Agentic-RAG
84
AIM-FinanceReasoning
87
AIM-Text-to-SQL
66
ARC-AGI-2
95
ARC-AGI-3
63
AutomationBench-AA
41
BrowseComp
92
Convex Coding Evals
85
CritPt
32
DeepSWE 1.1
74
FrontierCode
53
FrontierMath
98
GDPval
1542
GPQA Diamond
96
HealthBench Professional
70
RuneBench
7
SciCode
56
Terminal-Bench 2.1
90
Terminal-Bench 4.0
58

Cost$20.00
Latency3.84s
Context1M
TTFT3.84s
AA-LCR
80
AIM-Agentic-RAG
84
AIM-FinanceReasoning
90
AIM-Text-to-SQL
67
ARC-AGI-1
98
ARC-AGI-2
90
AutomationBench-AA
31
Convex Coding Evals
81
CritPt
31
DeepSWE 1.1
67
FrontierCode
51
FrontierMath
88
GDPval
1735
GPQA Diamond
94
HealthBench Professional
62
Humanity's Last Exam
59
LegalBench
89
RuneBench
6
SciCode
63
Tau3-Banking
47
Terminal-Bench 2.1
91
Terminal-Bench 4.0
56

Cost$0.23
Latency0.85s
Context1M
TTFT0.85s
AA-LCR
77
GDPval
1743
GPQA Diamond
92
Humanity's Last Exam
38
IFBench
81
SciCode
47
Tau3-Banking
45
Terminal-Bench 2.1
86

Cost$0.60
Latency-
Context200k
TTFT
CritPt
30
IFBench
74
LegalBench
61
Tau2-Bench Telecom
81
Terminal-Bench Hard
25

Cost$0.54
Latency43.98s
Context1M
TTFT43.98s
CritPt
27
GDPval
1673
SciCode
61

Cost$2.00
Latency5.26s
Context1M
TTFT5.26s
AA-LCR
81
DeepSWE 1.1
53
GDPval
1381
GPQA Diamond
90
HealthBench Professional
59
Humanity's Last Exam
46
LegalBench
85
ObviousBench
99
ReactBench
23
RuneBench
6
SciCode
58
Tau3-Banking
32
Terminal-Bench 2.1
78

Cost$20.00
Latency4.35s
Context1M
TTFT4.35s
AA-LCR
77
AIM-A-CODE-LLM Bench
69
AIM-Agentic-RAG
82
AIM-FinanceReasoning
90
AIM-Text-to-SQL
66
APEX-Agents-AA
59
ARC-AGI-1
99
ARC-AGI-2
89
ArxivMath
79
AutomationBench-AA
17
BrowseComp
87
Convex Coding Evals
84
CritPt
29
DeepSWE 1.1
70
Frontier-Bench
34
FrontierCode
54
FrontierMath
90
GDPval
1741
GPQA Diamond
93
Harvey LAB-AA
14
HealthBench Professional
63
HieroglyphBench
23
Humanity's Last Exam
56
IFBench
64
LegalBench
89
MMMU-Pro
81
ObviousBench
99
Opus Magnum Bench
22
ReactBench
47
RuneBench
6
SciCode
61
SimpleBench
82
Tau2-Bench Telecom
99
Tau3-Banking
38
Terminal-Bench 2.1
85
Terminal-Bench 4.0
42
Terminal-Bench Hard
63
Terminal-Bench-Science
21

Cost$10.00
Latency4.03s
Context1M
TTFT4.03s
AA-LCR
79
AIM-Agentic-RAG
85
AIM-FinanceReasoning
90
AIM-Text-to-SQL
64
ARC-AGI-1
98
ARC-AGI-2
90
ARC-AGI-3
30
ArxivMath
91
AutomationBench-AA
50
BioMysteryBench
79
BrowseComp
91
Convex Coding Evals
82
CritPt
29
DeepSWE 1.1
74
Frontier-Bench
44
FrontierCode
53
FrontierMath
73
GDPval
1708
GPQA Diamond
94
HealthBench Professional
60
Humanity's Last Exam
55
LegalBench
87
MMMU-Pro
85
ObviousBench
100
RuneBench
6
SciCode
56
Tau3-Banking
45
Terminal-Bench 2.1
89
Terminal-Bench 4.0
52
Terminal-Bench-Science
30

Cost$4.00
Latency1.36s
Context1M
TTFT1.36s
AIM-Agentic-RAG
82
AIM-Text-to-SQL
54
CritPt
31
GDPval
1487
SciCode
58
Page 1 of 13
Frontier Over Time
Intelligence Index by model release date
Cost vs Performance
Blended cost against Intelligence Index
Model × Benchmark
Top models on selected benchmarks

How the Index Is Built

Each benchmark is read as a placement rather than a raw score. Within one benchmark the best result sits at 100, the worst at 0, and every other model falls somewhere in between. The index is the weighted average of those placements over the benchmarks a model has actually run: a benchmark it was never run on is not counted as a zero, it is simply absent. Benchmarks do not count equally, and the weight each one carries is listed with it. A model needs results on at least two index benchmarks before it is ranked, so a single result places a model on one benchmark's line without giving it a standing of its own. Placements are used instead of raw accuracy because benchmarks differ sharply in difficulty and scale, and scores from different benchmarks cannot share an average. Weights: AIM-FinanceReasoning (10%) + AIM-Text-to-SQL (10%) + AIM-Agentic-RAG (10%) + SciCode (10%) + FrontierMath (10%) + Terminal-Bench 2.1 (10%) + CritPt (10%) + IFBench (10%) + MMMU-Pro (10%) + ARC-AGI-3 (10%).

Benchmarks We Used

Public
AA-LCRLong-context reasoning over multi-document inputs.
AIMultiple
AIM-A-CODE-LLM BenchAgentic coding across 10 software-dev tasks via CLI tool (~3,500 validation steps).
AIMultiple
AIM-Agentic-RAGMulti-database routing & SQL query generation across 5 databases (BIRD-SQL, 500 questions).
AIMultiple
AIM-FinanceReasoningComplex multi-step financial reasoning across 238 hard FinanceReasoning questions.
AIMultiple
AIM-HALC-BenchResistance to fabricating unmentioned metrics in long-context documents (204 traps, 14 transcripts).
AIMultiple
AIM-RELC-BenchLong-context numeric fact retrieval across document positions (100 items, 14 transcripts).
AIMultiple
AIM-Text-to-SQLNatural-language-to-SQL query generation accuracy across 35+ LLMs.
Public
APEX-Agents-AALong-horizon, cross-application tasks written by investment bankers, consultants and corporate lawyers.
Public
ARC-AGI-1Few-shot abstract reasoning over grid transformations; the first ARC Prize benchmark.
Public
ARC-AGI-2Novel visual reasoning puzzles testing generalization, not memorization.
Public
ARC-AGI-3Interactive reasoning: agents explore novel game environments, form goals on the fly and learn across steps. 100% equals human learning efficiency.
Public
ArxivMathRecent arXiv mathematics problems, run soon after publication to limit contamination.
Public
AutomationBench-AAEnd-to-end office automation tasks, scored on partial completion.
Public
BioMysteryBenchAnthropic's benchmark of whether AI agents can solve real bioinformatics mysteries from raw data.
Public
BrowseCompHard-to-find facts on the live web: short questions that need persistent browsing to answer.
Public
Convex Coding EvalsWriting Convex backends and client integrations without framework-specific guidelines in the prompt.
Public
CritPtResearch-level physics problems requiring multi-step derivation.
Public
Crosby RedlineBenchContract redlining tasks scored turn by turn against attorney edits.
Public
DeepSWE 1.1Long-horizon engineering tasks from live open-source repositories, graded on committed code in a clean environment.
Public
EnterpriseOps-Gym-AAStateful agentic planning and tool use across eight enterprise systems.
Public
Frontier-BenchAgentic software tasks on frontier repositories, scored by resolution rate.
Public
FrontierCodeWhether a change is mergeable, not just test-passing: regression safety, scope and maintainability judged by rubric.
Public
FrontierMathUnpublished, expert-authored research math problems (Tiers 1–4).
Public
GDPvalReal deliverables from 44 occupations across the nine largest sectors of US GDP, graded blind by experienced professionals against a human-made reference.
Public
GPQA DiamondGraduate-level Google-proof science questions that require deep reasoning.
Public
Harvey LAB-AAHarvey's legal agent benchmark, scored on all-pass task success.
Public
HealthBench ProfessionalPhysician-written conversations scored against physician-written rubrics, on the professional split.
Public
HieroglyphBenchReading Egyptian hieroglyphs from images, scored on sign accuracy.
Public
Humanity's Last ExamExpert-level, closed-book reasoning across 100+ academic subjects (2.5k questions).
Public
IFBenchPrecise instruction following on 58 unseen, verifiable output constraints.
Public
ITBench-AAReal-world IT automation scenarios across site reliability, FinOps and security operations.
Public
LegalBenchCrowd-sourced legal reasoning across 162 tasks from legal practitioners.
Public
LiveCodeBenchContamination-free, continuously updated competitive coding problems.
Public
MMMU-ProCollege-level multimodal understanding across disciplines, hardened against text-only shortcuts.
Public
ObviousBenchQuestions with obvious answers that models nonetheless get wrong; scored pass-cubed.
Public
Opus Magnum BenchDesigning working machines in the puzzle game Opus Magnum.
Public
ReactBenchBuilding and fixing React components, scored pass@1.
Public
RuneBenchReading and reasoning over a constructed rune script the model has never seen, scored on a log scale.
Public
SciCodeResearch-level scientific coding across physics, math, biology & chemistry (338 subproblems)
Public
SimpleBenchEveryday reasoning questions where humans reliably beat frontier models.
Public
Swe-BenchReal-world GitHub issues resolved end-to-end (full 2,294-instance set).
Public
Tau2-Bench TelecomTool-using customer support agents in a telecom domain, scored on task success.
Public
Tau3-BankingTool-using agents handling realistic banking customer service workflows.
Public
Terminal-Bench 2.1Agentic terminal tasks across software engineering, data and system administration.
Public
Terminal-Bench 4.0The current Terminal-Bench generation: agentic terminal tasks scored by resolution rate.
Public
Terminal-Bench HardThe hard subset of Terminal-Bench: multi-step software tasks solved in a real terminal.
Public
Terminal-Bench-ScienceTerminal tasks drawn from real scientific computing workflows.

Recent Updates

Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.

Anthropic

Claude Opus 5.5

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-6 Sol

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-6 Luna

New model added to the AIMultiple Intelligence Index.

Xiaomi

MiMo-V2.6-Pro

New model added to the AIMultiple Intelligence Index.

LLM Release TimelineThe pace of AI innovation
AI Labs
Alibaba Cloud
Alibaba Cloud
03 Aug 2026
Qwen3.7 Flash +1 moreQwen3.7 FlashQwen3.8 Max
14 Aug 2026
Qwen3.8 2.4T A95B +1 moreQwen3.8 2.4T A95BQwen3.8 27B
03 Sep 2026
Qwen3.8 Flash +1 moreQwen3.8 FlashQwen3.8 Max (0902)
23 Sep 2026
Qwen3.8 Max Prime +1 moreQwen3.8 Max PrimeQwen3.8 Omni Flash
Z AI
Z AI
26 Aug 2026
GLM 5.3 +1 moreGLM 5.3GLM 5.3 Flash
23 Sep 2026
GLM 5.3 FlashX +1 moreGLM 5.3 FlashXGLM 5.3 Prime
Anthropic
Anthropic
24 Jul 2026
Claude Opus 5
01 Sep 2026
Claude Fable 5.1
22 Sep 2026
Claude Opus 5.5
OpenAI
OpenAI
22 Sep 2026
GPT-6 Astra +2 moreGPT-6 AstraGPT-6 LunaGPT-6 Sol
Xiaomi
Xiaomi
21 Sep 2026
MiMo-V2.6-Pro
X
X
12 Aug 2026
Grok 4.6
21 Sep 2026
Grok 4.7
Meta
Meta
09 Aug 2026
Muse Glimmer 30B +1 moreMuse Glimmer 30BMuse Spark 1.2
02 Sep 2026
Muse Spark 1.2 Contributor +2 moreMuse Spark 1.2 ContributorMuse Spark 1.3Muse Spark 1.3 Contributor
Google
Google
13 Aug 2026
Gemini 3.7 Flash
02 Sep 2026
Gemini 3.8 Flash
Tencent
Tencent
28 Aug 2026
Hy4

Explore LLM Use Cases, Analyses & Benchmarks

LLM Parameters: GPT-5 High, Medium, Low and Minimal

LLM
Insight
Aug 21

Some LLMs, such as OpenAI’s GPT-5 family, come in different versions (e.g., GPT-5, GPT-5-mini, and GPT-5-nano) and with various parameter settings, including high, medium, low, and minimal. Below, we explore the differences between these model versions by gathering their benchmark performance and the costs to run the benchmarks. We used the GPT-5 family in our…

Read More
LLM
Feature Comparison
Aug 21

Top LLMOps Tools & Compare them to MLOPs

LLMOps platforms handle the operational side of running large language models: deployment, monitoring, evaluation, and cost management. We examined top LLMOps tools, their core features, pricing models, and how they differ from each other to help identify the best fit for various use cases. A breakdown of each metric is provided below: LLMOps platforms support…

LLM
Insight
Aug 19

50+ ChatGPT Use Cases with Real Life Examples

ChatGPT reached approximately 1 billion weekly active users in early 2026 roughly 10% of the world’s population.8 OpenAI surpassed $20 billion in annual revenue for 2025, confirmed by CFO Sarah Friar.9 The Anthropic Economic Index distinguishes two modes of use: augmentation, in which a human interacts with AI, and automation, in which AI completes tasks…

LLM
Insight
Aug 19

ChatGPT for Customer Service: Top 10 Use Cases

ChatGPT has moved from novelty to infrastructure in customer service. Companies are using it to cut response times, handle volume their teams can’t absorb, and reduce the cost of routine interactions. But results vary sharply depending on how it’s implemented. OpenAI launched GPT-5.6, a materially more capable model that is better at instruction-following, reasoning across…

LLM
Benchmark
Aug 16

Intelligence Density of 71 LLMs for Smarter & Denser Models

We tracked 71 LLMs released between February 2023 and May 2026 and collected 10 public benchmarks to measure intelligence density. We divided the capability score by the resource the model consumes (active parameters, training compute, and inference price). To calculate intelligence density, we executed the following steps: See methodology for the scoring approach, and per-resource…

LLM
Benchmark
Aug 12

LLM Latency Benchmark by Use Cases

We benchmarked 11 top large language models with a total of 1,320 requests, splitting reasoning and non-reasoning models, and measured first-token latency, per-token latency, and overall response time. You can find details on how we measured latency here. We report reasoning and non-reasoning models separately. Reasoning models spend several seconds thinking before the first visible…

LLM
Benchmark
Apr 15

LLM Quantization: BF16 vs FP8 vs INT4

We benchmarked Qwen3-32B at 4 precision levels (BF16, FP8, GPTQ-Int8, GPTQ-Int4) on a single NVIDIA H100 80GB GPU. Each configuration was evaluated on 2 benchmarks (~12.2K questions) covering knowledge and code generation, plus 2,000+ inference runs to measure throughput. Int4 is 2.7x faster than BF16 while losing less than 2 points on MMLU-Pro, but code…