The newest model in the AIMultiple Intelligence Index is Kimi K3.
LLM Benchmarks
One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations. See the methodology.
Leaderboard
The highest-scoring models across all benchmarks.
# | Model | Index | LegalBench | FinanceReasoning | Text-to-SQL | Agentic RAG | ARC-AGI-2 | Agentic LLM Benchmark | FrontierMath | Swe-Bench |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 Anthropic | 92 | 89 | 90 | 90 | 98 | - | 69 | - | - |
| 2 | Kimi K3 Moonshot AI | 88 | 86 | 88 | 77 | 93 | - | 73 | - | - |
| 3 | GPT-5.6 Sol OpenAI | 87 | 87 | 90 | 74 | 87 | 93 | 62 | - | - |
| 4 | Grok 4.5 X | 86 | 86 | 88 | 79 | 83 | - | 73 | - | - |
| 5 | GPT-5.5 OpenAI | 83 | 87 | - | - | - | 85 | 59 | - | - |
| 6 | GPT-5.6 Terra OpenAI | 80 | 85 | 87 | 71 | 91 | 84 | 61 | - | - |
| 7 | GPT-5.6 Sol Pro OpenAI | 80 | - | 91 | 79 | 96 | - | 54 | - | - |
| 8 | Claude Opus 4.6 Anthropic | 78 | - | 88 | 68 | 80 | 69 | 72 | - | - |
| 9 | Gemini 3 Pro Preview Google | 77 | 87 | 86 | 60 | 89 | 31 | - | - | - |
| 10 | Gemini 3.1 Pro Preview Google | 77 | 87 | 87 | 65 | 89 | 77 | 46 | 27 | - |
Page 1 of 7 | ||||||||||
How The Index is Built
The Intelligence Index normalizes each benchmark to a 0-100 scale and averages a model's relative standing across the benchmarks it was evaluated on. Each benchmark below feeds that score.
Recent Updates
Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.
Kimi K3
New model added to the AIMultiple Intelligence Index.
GPT-5.6 Sol
New model added to the AIMultiple Intelligence Index.
GPT-5.6 Terra
New model added to the AIMultiple Intelligence Index.
GPT-5.6 Sol Pro
New model added to the AIMultiple Intelligence Index.
Key AI model releases from leading providers over the past 41 days.
Explore LLM Use Cases, Analyses & Benchmarks
LLM Automation: Top 7 Tools & 8 Case Studies
LLM automation refers to shift to intelligent automation tools that leverage LLMs, including AI agents, fine-tuned LLMs and RAG models to automate and coordinate tasks. Explore what LLM automation is, its top real-life applications and major tools: Large language models in automation is a systematic approach that combines Natural Language Processing (NLP) with existing process…
Large Multimodal Models (LMMs) vs LLMs
We evaluated the performance of Large Multimodal Models (LMMs) in financial reasoning tasks using a carefully selected dataset. By analyzing a subset of high-quality financial samples, we assess the models’ capabilities in processing and reasoning with multimodal data in the financial domain. The methodology section provides detailed insights into the dataset and evaluation framework employed.…
Compare 9 Large Language Models in Healthcare
We benchmarked 9 LLMs using the MedQA dataset, a graduate-level clinical exam benchmark derived from USMLE questions. Each model answered the same multiple-choice clinical scenarios using a standardized prompt, enabling direct comparison of accuracy. We also recorded latency per question by dividing total runtime by the number of MedQA items completed. Benchmark methodology: This benchmark…
Intelligence Density of 71 LLMs for Smarter & Denser Models
We tracked 71 LLMs released between February 2023 and May 2026 and collected 10 public benchmarks to measure intelligence density. We divided the capability score by the resource the model consumes (active parameters, training compute, and inference price). To calculate intelligence density, we executed the following steps: See methodology for the scoring approach, and per-resource…
50+ ChatGPT Use Cases with Real Life Examples
ChatGPT reached approximately 1 billion weekly active users in early 2026 roughly 10% of the world’s population.50 OpenAI surpassed $20 billion in annual revenue for 2025, confirmed by CFO Sarah Friar.51 The Anthropic Economic Index distinguishes two modes of use: augmentation, in which a human interacts with AI, and automation, in which AI completes tasks…
Benchmark of 40+ LLMs in Finance: Claude Fable 5 & GPT-5.6 Sol
We evaluated LLMs on 238 hard questions from the FinanceReasoning benchmark (Tang et al.).67 This subset targets the most challenging financial-reasoning tasks, assessing complex, multi-step quantitative reasoning involving financial concepts and formulas. Our evaluation employed a custom prompt design and scoring criteria of accuracy and token consumption. For a detailed explanation of how these metrics…
AIM Enterprise: Agentic Enterprise Benchmark
Enterprises use LLMs everyday for their regular tasks. To find most cost efficient LLMs, we designed AIM Enterprise, an agentic enterprise benchmark, where we used 69 enterprise tasks in different categories. Each model delivered one file per task. Scores are relative: the judges rank every answer against the other answers to the same task, so…
LLM Scaling Laws: Analysis from AI Researchers
Large language models predict the next token based on patterns learned from text data. The term LLM scaling laws refers to empirical regularities that link model performance to the amount of compute, training data, and model parameters used during training. To understand how these relationships influence modern model design in practice, we reviewed findings from…
LLM Observability Tools: Weights & Biases, Langsmith
LLM applications have expanded from single-turn chats into multi-step agents that use tools, query databases, and coordinate with other models, making their behavior harder to interpret. LLM observability provides continuous visibility into these complex workflows, helping organizations monitor quality, detect failures, troubleshoot issues, and manage performance and costs. W&B Weave is Weights & Biases‘ LLM…
LLM Latency Benchmark by Use Cases in 2026
We benchmarked 11 top large language models with a total of 1,320 requests, splitting reasoning and non-reasoning models, and measured first-token latency, per-token latency, and overall response time. You can find details on how we measured latency here. We report reasoning and non-reasoning models separately. Reasoning models spend several seconds thinking before the first visible…