The newest model in the AIMultiple Intelligence Index is Kimi K3.
LLM Benchmarks
One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations. See the methodology.
Leaderboard
The highest-scoring models across all benchmarks.
# | Model | Index | LegalBench | FinanceReasoning | Text-to-SQL | Agentic RAG | ARC-AGI-2 | Agentic LLM Benchmark | FrontierMath | Swe-Bench |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 Anthropic | 92 | 89 | 90 | 90 | 98 | - | 69 | - | - |
| 2 | Kimi K3 Moonshot AI | 88 | 86 | 88 | 77 | 93 | - | 73 | - | - |
| 3 | GPT-5.6 Sol OpenAI | 87 | 87 | 90 | 74 | 87 | 93 | 62 | - | - |
| 4 | Grok 4.5 X | 86 | 86 | 88 | 79 | 83 | - | 73 | - | - |
| 5 | GPT-5.5 OpenAI | 83 | 87 | - | - | - | 85 | 59 | - | - |
| 6 | GPT-5.6 Terra OpenAI | 80 | 85 | 87 | 71 | 91 | 84 | 61 | - | - |
| 7 | GPT-5.6 Sol Pro OpenAI | 80 | - | 91 | 79 | 96 | - | 54 | - | - |
| 8 | Claude Opus 4.6 Anthropic | 78 | - | 88 | 68 | 80 | 69 | 72 | - | - |
| 9 | Gemini 3 Pro Preview Google | 77 | 87 | 86 | 60 | 89 | 31 | - | - | - |
| 10 | Gemini 3.1 Pro Preview Google | 77 | 87 | 87 | 65 | 89 | 77 | 46 | 27 | - |
Page 1 of 7 | ||||||||||
How The Index is Built
The Intelligence Index normalizes each benchmark to a 0-100 scale and averages a model's relative standing across the benchmarks it was evaluated on. Each benchmark below feeds that score.
Recent Updates
Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.
Kimi K3
New model added to the AIMultiple Intelligence Index.
GPT-5.6 Sol
New model added to the AIMultiple Intelligence Index.
GPT-5.6 Terra
New model added to the AIMultiple Intelligence Index.
GPT-5.6 Sol Pro
New model added to the AIMultiple Intelligence Index.
Explore LLM Use Cases, Analyses & Benchmarks
50+ ChatGPT Use Cases with Real Life Examples
ChatGPT reached approximately 1 billion weekly active users in early 2026 roughly 10% of the world’s population.1 OpenAI surpassed $20 billion in annual revenue for 2025, confirmed by CFO Sarah Friar.2 The Anthropic Economic Index distinguishes two modes of use: augmentation, in which a human interacts with AI, and automation, in which AI completes tasks…
Benchmark of 40+ LLMs in Finance: Claude Fable 5 & GPT-5.6 Sol
We evaluated LLMs on 238 hard questions from the FinanceReasoning benchmark (Tang et al.).18 This subset targets the most challenging financial-reasoning tasks, assessing complex, multi-step quantitative reasoning involving financial concepts and formulas. Our evaluation employed a custom prompt design and scoring criteria of accuracy and token consumption. For a detailed explanation of how these metrics…
AIM Enterprise: Agentic Enterprise Benchmark
Enterprises use LLMs everyday for their regular tasks. To find most cost efficient LLMs, we designed AIM Enterprise, an agentic enterprise benchmark, where we used 69 enterprise tasks in different categories. Each model delivered one file per task. Scores are relative: the judges rank every answer against the other answers to the same task, so…
LLM Scaling Laws: Analysis from AI Researchers
Large language models predict the next token based on patterns learned from text data. The term LLM scaling laws refers to empirical regularities that link model performance to the amount of compute, training data, and model parameters used during training. To understand how these relationships influence modern model design in practice, we reviewed findings from…
LLM Observability Tools: Weights & Biases, Langsmith
LLM applications have expanded from single-turn chats into multi-step agents that use tools, query databases, and coordinate with other models, making their behavior harder to interpret. LLM observability provides continuous visibility into these complex workflows, helping organizations monitor quality, detect failures, troubleshoot issues, and manage performance and costs. W&B Weave is Weights & Biases‘ LLM…
Large Multimodal Models (LMMs) vs LLMs
We evaluated the performance of Large Multimodal Models (LMMs) in financial reasoning tasks using a carefully selected dataset. By analyzing a subset of high-quality financial samples, we assess the models’ capabilities in processing and reasoning with multimodal data in the financial domain. The methodology section provides detailed insights into the dataset and evaluation framework employed.…
ChatGPT for Customer Service: Top 10 Use Cases
ChatGPT has moved from novelty to infrastructure in customer service. Companies are using it to cut response times, handle volume their teams can’t absorb, and reduce the cost of routine interactions. But results vary sharply depending on how it’s implemented. OpenAI launched GPT-5.6, a materially more capable model that is better at instruction-following, reasoning across…
LLM Latency Benchmark by Use Cases in 2026
We benchmarked 11 top large language models with a total of 1,320 requests, splitting reasoning and non-reasoning models, and measured first-token latency, per-token latency, and overall response time. You can find details on how we measured latency here. We report reasoning and non-reasoning models separately. Reasoning models spend several seconds thinking before the first visible…
Top LLMOps Tools & Compare them to MLOPs
LLMOps platforms handle the operational side of running large language models: deployment, monitoring, evaluation, and cost management. We examined top LLMOps tools, their core features, pricing models, and how they differ from each other to help identify the best fit for various use cases. A breakdown of each metric is provided below: LLMOps platforms support…
LLM Pricing: Top 15+ Providers Compared
LLM pricing spans three orders of magnitude: the cheapest commodity models cost under $0.20 per million tokens, while frontier reasoning tiers launched as high as $262.50. The chart below tracks how launch prices moved: each model sits at its launch date with its launch list price per million tokens, blended at a 3:1 input-to-output ratio,…