The newest model in the AIMultiple Intelligence Index is Claude Opus 5.5.
LLM Benchmarks
One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations. How the Index Is Built
Leaderboard
The highest-scoring models across all benchmarks.
# | Model | Index | AA-LCR | AIM-A-CODE-LLM Bench | AIM-Agentic-RAG | AIM-FinanceReasoning | AIM-HALC-Bench | AIM-RELC-Bench | AIM-Text-to-SQL | APEX-Agents-AA | ARC-AGI-1 | ARC-AGI-2 | ARC-AGI-3 | ArxivMath | AutomationBench-AA | BioMysteryBench | BrowseComp | Convex Coding Evals | CritPt | Crosby RedlineBench | DeepSWE 1.1 | EnterpriseOps-Gym-AA | Frontier-Bench | FrontierCode | FrontierMath | GDPval | GPQA Diamond | Harvey LAB-AA | HealthBench Professional | HieroglyphBench | Humanity's Last Exam | IFBench | ITBench-AA | LegalBench | LiveCodeBench | MMMU-Pro | ObviousBench | Opus Magnum Bench | ReactBench | RuneBench | SciCode | SimpleBench | Swe-Bench | Tau2-Bench Telecom | Tau3-Banking | Terminal-Bench 2.1 | Terminal-Bench 4.0 | Terminal-Bench Hard | Terminal-Bench-Science |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 Anthropic | 100 | - | - | 90 | 92 | - | - | 88 | - | - | - | - | - | - | - | - | - | 32 | - | - | - | - | - | - | 1846 | - | - | - | - | - | - | - | - | - | - | - | - | - | - | 67 | - | - | - | - | - | - | - | - |
| 2 | GPT-6 Astra OpenAI | 92 | - | - | 84 | 87 | - | - | 66 | - | - | 95 | 63 | - | 41 | - | 92 | 85 | 32 | - | 74 | - | - | 53 | 98 | 1542 | 96 | - | 70 | - | - | - | - | - | - | - | - | - | - | 7 | 56 | - | - | - | - | 90 | 58 | - | - |
| 3 | Claude Fable 5.1 Anthropic | 91 | 80 | - | 84 | 90 | - | - | 67 | - | 98 | 90 | - | - | 31 | - | - | 81 | 31 | - | 67 | - | - | 51 | 88 | 1735 | 94 | - | 62 | - | 59 | - | - | 89 | - | - | - | - | - | 6 | 63 | - | - | - | 47 | 91 | 56 | - | - |
| 4 | Qwen3.8 Flash Alibaba Cloud | 87 | 77 | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | 1743 | 92 | - | - | - | 38 | 81 | - | - | - | - | - | - | - | - | 47 | - | - | - | 45 | 86 | - | - | - |
| 5 | Command A+ Cohere | 87 | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | 30 | - | - | - | - | - | - | - | - | - | - | - | - | 74 | - | 61 | - | - | - | - | - | - | - | - | - | 81 | - | - | - | 25 | - |
| 6 | MiMo-V2.6-Pro Xiaomi | 87 | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | 27 | - | - | - | - | - | - | 1673 | - | - | - | - | - | - | - | - | - | - | - | - | - | - | 61 | - | - | - | - | - | - | - | - |
| 7 | Muse Spark 1.1 Meta | 86 | 81 | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | 53 | - | - | - | - | 1381 | 90 | - | 59 | - | 46 | - | - | 85 | - | - | 99 | - | 23 | 6 | 58 | - | - | - | 32 | 78 | - | - | - |
| 8 | Claude Fable 5 Anthropic | 85 | 77 | 69 | 82 | 90 | - | - | 66 | 59 | 99 | 89 | - | 79 | 17 | - | 87 | 84 | 29 | - | 70 | - | 34 | 54 | 90 | 1741 | 93 | 14 | 63 | 23 | 56 | 64 | - | 89 | - | 81 | 99 | 22 | 47 | 6 | 61 | 82 | - | 99 | 38 | 85 | 42 | 63 | 21 |
| 9 | Claude Opus 5 Anthropic | 83 | 79 | - | 85 | 90 | - | - | 64 | - | 98 | 90 | 30 | 91 | 50 | 79 | 91 | 82 | 29 | - | 74 | - | 44 | 53 | 73 | 1708 | 94 | - | 60 | - | 55 | - | - | 87 | - | 85 | 100 | - | - | 6 | 56 | - | - | - | 45 | 89 | 52 | - | 30 |
| 10 | GPT-6 Sol OpenAI | 79 | - | - | 82 | - | - | - | 54 | - | - | - | - | - | - | - | - | - | 31 | - | - | - | - | - | - | 1487 | - | - | - | - | - | - | - | - | - | - | - | - | - | - | 58 | - | - | - | - | - | - | - | - |
How the Index Is Built
Each benchmark is read as a placement rather than a raw score. Within one benchmark the best result sits at 100, the worst at 0, and every other model falls somewhere in between. The index is the weighted average of those placements over the benchmarks a model has actually run: a benchmark it was never run on is not counted as a zero, it is simply absent. Benchmarks do not count equally, and the weight each one carries is listed with it. A model needs results on at least two index benchmarks before it is ranked, so a single result places a model on one benchmark's line without giving it a standing of its own. Placements are used instead of raw accuracy because benchmarks differ sharply in difficulty and scale, and scores from different benchmarks cannot share an average. Weights: AIM-FinanceReasoning (10%) + AIM-Text-to-SQL (10%) + AIM-Agentic-RAG (10%) + SciCode (10%) + FrontierMath (10%) + Terminal-Bench 2.1 (10%) + CritPt (10%) + IFBench (10%) + MMMU-Pro (10%) + ARC-AGI-3 (10%).
Benchmarks We Used
Recent Updates
Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.
Claude Opus 5.5
New model added to the AIMultiple Intelligence Index.
GPT-6 Sol
New model added to the AIMultiple Intelligence Index.
GPT-6 Luna
New model added to the AIMultiple Intelligence Index.
MiMo-V2.6-Pro
New model added to the AIMultiple Intelligence Index.
Explore LLM Use Cases, Analyses & Benchmarks
Top LLMOps Tools & Compare them to MLOPs
LLMOps platforms handle the operational side of running large language models: deployment, monitoring, evaluation, and cost management. We examined top LLMOps tools, their core features, pricing models, and how they differ from each other to help identify the best fit for various use cases. A breakdown of each metric is provided below: LLMOps platforms support…
50+ ChatGPT Use Cases with Real Life Examples
ChatGPT reached approximately 1 billion weekly active users in early 2026 roughly 10% of the world’s population.8 OpenAI surpassed $20 billion in annual revenue for 2025, confirmed by CFO Sarah Friar.9 The Anthropic Economic Index distinguishes two modes of use: augmentation, in which a human interacts with AI, and automation, in which AI completes tasks…
ChatGPT for Customer Service: Top 10 Use Cases
ChatGPT has moved from novelty to infrastructure in customer service. Companies are using it to cut response times, handle volume their teams can’t absorb, and reduce the cost of routine interactions. But results vary sharply depending on how it’s implemented. OpenAI launched GPT-5.6, a materially more capable model that is better at instruction-following, reasoning across…
Intelligence Density of 71 LLMs for Smarter & Denser Models
We tracked 71 LLMs released between February 2023 and May 2026 and collected 10 public benchmarks to measure intelligence density. We divided the capability score by the resource the model consumes (active parameters, training compute, and inference price). To calculate intelligence density, we executed the following steps: See methodology for the scoring approach, and per-resource…
LLM Latency Benchmark by Use Cases
We benchmarked 11 top large language models with a total of 1,320 requests, splitting reasoning and non-reasoning models, and measured first-token latency, per-token latency, and overall response time. You can find details on how we measured latency here. We report reasoning and non-reasoning models separately. Reasoning models spend several seconds thinking before the first visible…
LLM Quantization: BF16 vs FP8 vs INT4
We benchmarked Qwen3-32B at 4 precision levels (BF16, FP8, GPTQ-Int8, GPTQ-Int4) on a single NVIDIA H100 80GB GPU. Each configuration was evaluated on 2 benchmarks (~12.2K questions) covering knowledge and code generation, plus 2,000+ inference runs to measure throughput. Int4 is 2.7x faster than BF16 while losing less than 2 points on MMLU-Pro, but code…