Premium
Services
Premium

LLM Benchmarks

One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations. How the Index Is Built

TOP MODEL
Claude Opus 5.5
Index 100
Best value
Qwen3.8 Flash
$0.23 / 1M
Fastest
MiniMax M2.7
0.15s TTFT
AIMultiple Intelligence Index

Leaderboard

The highest-scoring models across all benchmarks.

Filter & Sort
#
Model
Index
AA-LCR
AIM-A-CODE-LLM Bench
AIM-Agentic-RAG
AIM-FinanceReasoning
AIM-HALC-Bench
AIM-RELC-Bench
AIM-Text-to-SQL
APEX-Agents-AA
ARC-AGI-1
ARC-AGI-2
ARC-AGI-3
ArxivMath
AutomationBench-AA
BioMysteryBench
BrowseComp
Convex Coding Evals
CritPt
Crosby RedlineBench
DeepSWE 1.1
EnterpriseOps-Gym-AA
Frontier-Bench
FrontierCode
FrontierMath
GDPval
GPQA Diamond
Harvey LAB-AA
HealthBench Professional
HieroglyphBench
Humanity's Last Exam
IFBench
ITBench-AA
LegalBench
LiveCodeBench
MMMU-Pro
ObviousBench
Opus Magnum Bench
ReactBench
RuneBench
SciCode
SimpleBench
Swe-Bench
Tau2-Bench Telecom
Tau3-Banking
Terminal-Bench 2.1
Terminal-Bench 4.0
Terminal-Bench Hard
Terminal-Bench-Science
1
Claude Opus 5.5
Claude Opus 5.5
Anthropic
100
--9092--88---------32------1846--------------67--------
2
GPT-6 Astra
GPT-6 Astra
OpenAI
92
--8487--66--9563-41-928532-74--5398154296-70----------756----9058--
3
Claude Fable 5.1
Claude Fable 5.1
Anthropic
91
80-8490--67-9890--31--8131-67--5188173594-62-59--89-----663---479156--
4
Qwen3.8 Flash
Qwen3.8 Flash
Alibaba Cloud
87
77----------------------174392---3881--------47---4586---
5
Command A+
Command A+
Cohere
87
----------------30------------74-61---------81---25-
6
MiMo-V2.6-Pro
MiMo-V2.6-Pro
Xiaomi
87
----------------27------1673--------------61--------
7
Muse Spark 1.1
Muse Spark 1.1
Meta
86
81-----------------53----138190-59-46--85--99-23658---3278---
8
Claude Fable 5
Claude Fable 5
Anthropic
85
77698290--66599989-7917-878429-70-3454901741931463235664-89-8199224766182-993885426321
9
Claude Opus 5
Claude Opus 5
Anthropic
83
79-8590--64-989030915079918229-74-445373170894-60-55--87-85100--656---458952-30
10
GPT-6 Sol
GPT-6 Sol
OpenAI
79
--82---54---------31------1487--------------58--------
Page 1 of 13

Cost$8.00
Latency2.81s
Context1M
TTFT2.81s
AIM-Agentic-RAG
90
AIM-FinanceReasoning
92
AIM-Text-to-SQL
88
CritPt
32
GDPval
1846
SciCode
67

Cost$20.00
Latency6.31s
Context1M
TTFT6.31s
AIM-Agentic-RAG
84
AIM-FinanceReasoning
87
AIM-Text-to-SQL
66
ARC-AGI-2
95
ARC-AGI-3
63
AutomationBench-AA
41
BrowseComp
92
Convex Coding Evals
85
CritPt
32
DeepSWE 1.1
74
FrontierCode
53
FrontierMath
98
GDPval
1542
GPQA Diamond
96
HealthBench Professional
70
RuneBench
7
SciCode
56
Terminal-Bench 2.1
90
Terminal-Bench 4.0
58

Cost$20.00
Latency3.84s
Context1M
TTFT3.84s
AA-LCR
80
AIM-Agentic-RAG
84
AIM-FinanceReasoning
90
AIM-Text-to-SQL
67
ARC-AGI-1
98
ARC-AGI-2
90
AutomationBench-AA
31
Convex Coding Evals
81
CritPt
31
DeepSWE 1.1
67
FrontierCode
51
FrontierMath
88
GDPval
1735
GPQA Diamond
94
HealthBench Professional
62
Humanity's Last Exam
59
LegalBench
89
RuneBench
6
SciCode
63
Tau3-Banking
47
Terminal-Bench 2.1
91
Terminal-Bench 4.0
56

Cost$0.23
Latency0.85s
Context1M
TTFT0.85s
AA-LCR
77
GDPval
1743
GPQA Diamond
92
Humanity's Last Exam
38
IFBench
81
SciCode
47
Tau3-Banking
45
Terminal-Bench 2.1
86

Cost$0.60
Latency-
Context200k
TTFT
CritPt
30
IFBench
74
LegalBench
61
Tau2-Bench Telecom
81
Terminal-Bench Hard
25

Cost$0.54
Latency43.98s
Context1M
TTFT43.98s
CritPt
27
GDPval
1673
SciCode
61

Cost$2.00
Latency5.26s
Context1M
TTFT5.26s
AA-LCR
81
DeepSWE 1.1
53
GDPval
1381
GPQA Diamond
90
HealthBench Professional
59
Humanity's Last Exam
46
LegalBench
85
ObviousBench
99
ReactBench
23
RuneBench
6
SciCode
58
Tau3-Banking
32
Terminal-Bench 2.1
78

Cost$20.00
Latency4.35s
Context1M
TTFT4.35s
AA-LCR
77
AIM-A-CODE-LLM Bench
69
AIM-Agentic-RAG
82
AIM-FinanceReasoning
90
AIM-Text-to-SQL
66
APEX-Agents-AA
59
ARC-AGI-1
99
ARC-AGI-2
89
ArxivMath
79
AutomationBench-AA
17
BrowseComp
87
Convex Coding Evals
84
CritPt
29
DeepSWE 1.1
70
Frontier-Bench
34
FrontierCode
54
FrontierMath
90
GDPval
1741
GPQA Diamond
93
Harvey LAB-AA
14
HealthBench Professional
63
HieroglyphBench
23
Humanity's Last Exam
56
IFBench
64
LegalBench
89
MMMU-Pro
81
ObviousBench
99
Opus Magnum Bench
22
ReactBench
47
RuneBench
6
SciCode
61
SimpleBench
82
Tau2-Bench Telecom
99
Tau3-Banking
38
Terminal-Bench 2.1
85
Terminal-Bench 4.0
42
Terminal-Bench Hard
63
Terminal-Bench-Science
21

Cost$10.00
Latency4.03s
Context1M
TTFT4.03s
AA-LCR
79
AIM-Agentic-RAG
85
AIM-FinanceReasoning
90
AIM-Text-to-SQL
64
ARC-AGI-1
98
ARC-AGI-2
90
ARC-AGI-3
30
ArxivMath
91
AutomationBench-AA
50
BioMysteryBench
79
BrowseComp
91
Convex Coding Evals
82
CritPt
29
DeepSWE 1.1
74
Frontier-Bench
44
FrontierCode
53
FrontierMath
73
GDPval
1708
GPQA Diamond
94
HealthBench Professional
60
Humanity's Last Exam
55
LegalBench
87
MMMU-Pro
85
ObviousBench
100
RuneBench
6
SciCode
56
Tau3-Banking
45
Terminal-Bench 2.1
89
Terminal-Bench 4.0
52
Terminal-Bench-Science
30

Cost$4.00
Latency1.36s
Context1M
TTFT1.36s
AIM-Agentic-RAG
82
AIM-Text-to-SQL
54
CritPt
31
GDPval
1487
SciCode
58
Page 1 of 13
Frontier Over Time
Intelligence Index by model release date
Cost vs Performance
Blended cost against Intelligence Index
Model × Benchmark
Top models on selected benchmarks

How the Index Is Built

Each benchmark is read as a placement rather than a raw score. Within one benchmark the best result sits at 100, the worst at 0, and every other model falls somewhere in between. The index is the weighted average of those placements over the benchmarks a model has actually run: a benchmark it was never run on is not counted as a zero, it is simply absent. Benchmarks do not count equally, and the weight each one carries is listed with it. A model needs results on at least two index benchmarks before it is ranked, so a single result places a model on one benchmark's line without giving it a standing of its own. Placements are used instead of raw accuracy because benchmarks differ sharply in difficulty and scale, and scores from different benchmarks cannot share an average. Weights: AIM-FinanceReasoning (10%) + AIM-Text-to-SQL (10%) + AIM-Agentic-RAG (10%) + SciCode (10%) + FrontierMath (10%) + Terminal-Bench 2.1 (10%) + CritPt (10%) + IFBench (10%) + MMMU-Pro (10%) + ARC-AGI-3 (10%).

Benchmarks We Used

Public
AA-LCRLong-context reasoning over multi-document inputs.
AIMultiple
AIM-A-CODE-LLM BenchAgentic coding across 10 software-dev tasks via CLI tool (~3,500 validation steps).
AIMultiple
AIM-Agentic-RAGMulti-database routing & SQL query generation across 5 databases (BIRD-SQL, 500 questions).
AIMultiple
AIM-FinanceReasoningComplex multi-step financial reasoning across 238 hard FinanceReasoning questions.
AIMultiple
AIM-HALC-BenchResistance to fabricating unmentioned metrics in long-context documents (204 traps, 14 transcripts).
AIMultiple
AIM-RELC-BenchLong-context numeric fact retrieval across document positions (100 items, 14 transcripts).
AIMultiple
AIM-Text-to-SQLNatural-language-to-SQL query generation accuracy across 35+ LLMs.
Public
APEX-Agents-AALong-horizon, cross-application tasks written by investment bankers, consultants and corporate lawyers.
Public
ARC-AGI-1Few-shot abstract reasoning over grid transformations; the first ARC Prize benchmark.
Public
ARC-AGI-2Novel visual reasoning puzzles testing generalization, not memorization.
Public
ARC-AGI-3Interactive reasoning: agents explore novel game environments, form goals on the fly and learn across steps. 100% equals human learning efficiency.
Public
ArxivMathRecent arXiv mathematics problems, run soon after publication to limit contamination.
Public
AutomationBench-AAEnd-to-end office automation tasks, scored on partial completion.
Public
BioMysteryBenchAnthropic's benchmark of whether AI agents can solve real bioinformatics mysteries from raw data.
Public
BrowseCompHard-to-find facts on the live web: short questions that need persistent browsing to answer.
Public
Convex Coding EvalsWriting Convex backends and client integrations without framework-specific guidelines in the prompt.
Public
CritPtResearch-level physics problems requiring multi-step derivation.
Public
Crosby RedlineBenchContract redlining tasks scored turn by turn against attorney edits.
Public
DeepSWE 1.1Long-horizon engineering tasks from live open-source repositories, graded on committed code in a clean environment.
Public
EnterpriseOps-Gym-AAStateful agentic planning and tool use across eight enterprise systems.
Public
Frontier-BenchAgentic software tasks on frontier repositories, scored by resolution rate.
Public
FrontierCodeWhether a change is mergeable, not just test-passing: regression safety, scope and maintainability judged by rubric.
Public
FrontierMathUnpublished, expert-authored research math problems (Tiers 1–4).
Public
GDPvalReal deliverables from 44 occupations across the nine largest sectors of US GDP, graded blind by experienced professionals against a human-made reference.
Public
GPQA DiamondGraduate-level Google-proof science questions that require deep reasoning.
Public
Harvey LAB-AAHarvey's legal agent benchmark, scored on all-pass task success.
Public
HealthBench ProfessionalPhysician-written conversations scored against physician-written rubrics, on the professional split.
Public
HieroglyphBenchReading Egyptian hieroglyphs from images, scored on sign accuracy.
Public
Humanity's Last ExamExpert-level, closed-book reasoning across 100+ academic subjects (2.5k questions).
Public
IFBenchPrecise instruction following on 58 unseen, verifiable output constraints.
Public
ITBench-AAReal-world IT automation scenarios across site reliability, FinOps and security operations.
Public
LegalBenchCrowd-sourced legal reasoning across 162 tasks from legal practitioners.
Public
LiveCodeBenchContamination-free, continuously updated competitive coding problems.
Public
MMMU-ProCollege-level multimodal understanding across disciplines, hardened against text-only shortcuts.
Public
ObviousBenchQuestions with obvious answers that models nonetheless get wrong; scored pass-cubed.
Public
Opus Magnum BenchDesigning working machines in the puzzle game Opus Magnum.
Public
ReactBenchBuilding and fixing React components, scored pass@1.
Public
RuneBenchReading and reasoning over a constructed rune script the model has never seen, scored on a log scale.
Public
SciCodeResearch-level scientific coding across physics, math, biology & chemistry (338 subproblems)
Public
SimpleBenchEveryday reasoning questions where humans reliably beat frontier models.
Public
Swe-BenchReal-world GitHub issues resolved end-to-end (full 2,294-instance set).
Public
Tau2-Bench TelecomTool-using customer support agents in a telecom domain, scored on task success.
Public
Tau3-BankingTool-using agents handling realistic banking customer service workflows.
Public
Terminal-Bench 2.1Agentic terminal tasks across software engineering, data and system administration.
Public
Terminal-Bench 4.0The current Terminal-Bench generation: agentic terminal tasks scored by resolution rate.
Public
Terminal-Bench HardThe hard subset of Terminal-Bench: multi-step software tasks solved in a real terminal.
Public
Terminal-Bench-ScienceTerminal tasks drawn from real scientific computing workflows.

Recent Updates

Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.

Anthropic

Claude Opus 5.5

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-6 Sol

New model added to the AIMultiple Intelligence Index.

OpenAI

GPT-6 Luna

New model added to the AIMultiple Intelligence Index.

Xiaomi

MiMo-V2.6-Pro

New model added to the AIMultiple Intelligence Index.

LLM Release TimelineThe pace of AI innovation
AI Labs
Alibaba Cloud
Alibaba Cloud
03 Aug 2026
Qwen3.7 Flash +1 moreQwen3.7 FlashQwen3.8 Max
14 Aug 2026
Qwen3.8 2.4T A95B +1 moreQwen3.8 2.4T A95BQwen3.8 27B
03 Sep 2026
Qwen3.8 Flash +1 moreQwen3.8 FlashQwen3.8 Max (0902)
23 Sep 2026
Qwen3.8 Max Prime +1 moreQwen3.8 Max PrimeQwen3.8 Omni Flash
Z AI
Z AI
26 Aug 2026
GLM 5.3 +1 moreGLM 5.3GLM 5.3 Flash
23 Sep 2026
GLM 5.3 FlashX +1 moreGLM 5.3 FlashXGLM 5.3 Prime
Anthropic
Anthropic
24 Jul 2026
Claude Opus 5
01 Sep 2026
Claude Fable 5.1
22 Sep 2026
Claude Opus 5.5
OpenAI
OpenAI
22 Sep 2026
GPT-6 Astra +2 moreGPT-6 AstraGPT-6 LunaGPT-6 Sol
Xiaomi
Xiaomi
21 Sep 2026
MiMo-V2.6-Pro
X
X
12 Aug 2026
Grok 4.6
21 Sep 2026
Grok 4.7
Meta
Meta
09 Aug 2026
Muse Glimmer 30B +1 moreMuse Glimmer 30BMuse Spark 1.2
02 Sep 2026
Muse Spark 1.2 Contributor +2 moreMuse Spark 1.2 ContributorMuse Spark 1.3Muse Spark 1.3 Contributor
Google
Google
13 Aug 2026
Gemini 3.7 Flash
02 Sep 2026
Gemini 3.8 Flash
Tencent
Tencent
28 Aug 2026
Hy4

Explore LLM Use Cases, Analyses & Benchmarks

Text-to-SQL Benchmark: SQL Accuracy Across 40+ LLMs

LLM
Benchmark
Sep 25

SQL accuracy is the percentage of scored queries that return the reference result. Incorrect routes and references flagged as broken are excluded. A reference query is the SQL supplied as the expected answer. Each model reaches a different set of SQL questions because scoring depends on its database choices. Differences in these subsets and execution…

Read More
LLM
Insight
Sep 25

LLM Market Share: Compare Usage & Adoption

We analyzed LLM market share by combining usage-based data and web visit estimates to show how demand for large language models is distributed across AI labs and AI applications. Read the methodology to see how we measured and calculated these results. The United States dominated web visits across all four months, consistently accounting for 85–95%.…

LLM
Benchmark
Sep 24

Benchmark of 40+ LLMs in Finance: Claude Opus 5.5 & GPT-6 Astra

We evaluated LLMs on 238 hard questions from the FinanceReasoning benchmark (Tang et al.).6 This subset targets the most challenging financial-reasoning tasks, assessing complex, multi-step quantitative reasoning involving financial concepts and formulas. Our evaluation employed a custom prompt design and scoring criteria of accuracy and token consumption. For a detailed explanation of how these metrics…

LLM
Benchmark
Sep 21

Compare Multimodal AI Models on Visual Reasoning

We benchmarked 15 leading multimodal AI models on visual reasoning using 200 visual-based questions. The evaluation consisted of two tracks: 100 chart understanding questions testing data visualization interpretation, and 100 visual logic questions assessing pattern recognition and spatial reasoning. Each question was run 5 times to ensure consistent and reliable results. See our benchmark methodology…

LLM
Insight
Sep 21

Large Multimodal Models (LMMs) vs LLMs

Evaluate LLMs and LMMs by comparing their benchmark scores and real-world latency by clicking the model’s name in the table below. You can also weigh their input and output pricing to judge overall efficiency and value. *Audio is native on the E2B, E4B and 12B models only. Available in five sizes: E2B, E4B, 12B, 26B…

LLM
Insight
Sep 21

LLM VRAM Calculator for Self-Hosting

Self-hosting an LLM means running inference on hardware the operator controls rather than via a third-party API, which changes the cost, data control, and privacy profile. Whether a model runs at all depends on memory. The calculator estimates the VRAM or unified memory a model needs to run locally, based on the model, its precision,…

LLM
Benchmark
Sep 18

Agentic IT: Can AI Agents Design a Benchmark

We tested 16 models on a benchmark design in text-to-SQL and tool calling. Each model built one benchmark per topic, for a total of 32 submissions. No submission passed every rubric criterion. The agents could build and run tests, but none demonstrated both a blank-answer test and a correct-answer test of their own scorer. Text-to-SQL…

LLM
Benchmark
Sep 17

AIM Enterprise: Agentic Enterprise Benchmark

Enterprises use LLMs every day for their regular tasks. To find the most cost-efficient LLMs, we designed AIM Enterprise, an agentic enterprise benchmark, where we used 69 real enterprise tasks across strategy, marketing, HR, sales, and operations. Two judge models scored every file that passed the checks. On 32.7% of the individual scores the two…

LLM
Insight
Sep 17

The Future of Large Language Models

See the future of large language models by delving into promising approaches, such as self-training, fact-checking, and sparse expertise that could address LLM limitations. Success rate comparison of LLM’s Claude Sonnet 4.6 led the benchmark with an overall score of 0.748, with base and thinking variants tied to three decimal places. Claude Opus 4.8 (0.702),…

LLM
Benchmark
Sep 15

Large Language Models in Cybersecurity

We evaluated 7 large language models across 9 cybersecurity domains using SecBench, a large-scale and multi-format benchmark for security tasks. We tested each model on 44,823 multiple-choice questions (MCQs) and 3,087 short-answer questions (SAQs), covering data security, identity & access management, network security, vulnerability management, and cloud security. MCQs (Multiple-Choice Questions) benchmarking: SAQs (Short Answer…

LLM
Insight
Sep 15

LLM Scaling Laws: Analysis from AI Researchers

Large language models predict the next token based on patterns learned from text data. The term LLM scaling laws refers to empirical regularities that link model performance to the amount of compute, training data, and model parameters used during training. To understand how these relationships influence modern model design in practice, we reviewed findings from…