The newest model in the AIMultiple Intelligence Index is Claude Opus 5.5.
LLM Benchmarks
One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations. How the Index Is Built
Leaderboard
The highest-scoring models across all benchmarks.
# | Model | Index | AIM-Agentic-RAG | AIM-AI-Bias | AIM-FinanceReasoning | AIM-Text-to-SQL | ARC-AGI-3 | CritPt | FrontierMath | IFBench | MMMU-Pro | SciCode | Terminal-Bench 2.1 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 Anthropic | 79 | 90 | 94 | 92 | 88 | - | 32 | - | - | 88 | 67 | - |
| 2 | GPT-6 Astra OpenAI | 76 | 84 | 98 | 87 | 66 | 63 | 32 | 98 | - | 87 | 56 | 90 |
| 3 | Claude Fable 5.1 Anthropic | 73 | 84 | - | 90 | 67 | - | 31 | 88 | - | - | 63 | 91 |
| 4 | Qwen3.8 Flash Alibaba Cloud | 71 | - | - | - | - | - | - | - | 81 | - | 47 | 86 |
| 5 | Qwen3.8 Max Alibaba Cloud | 70 | 83 | 93 | - | 51 | - | - | 46 | - | 82 | 53 | 81 |
| 6 | Claude Opus 5 Anthropic | 68 | 85 | 81 | 90 | 64 | 30 | 29 | 73 | - | 85 | 56 | 89 |
| 7 | Muse Spark 1.1 Meta | 68 | - | - | - | - | - | - | - | - | - | 58 | 78 |
| 8 | GPT-5.6 Sol OpenAI | 67 | 77 | 89 | 90 | 54 | 8 | 32 | 83 | 73 | 83 | 57 | 90 |
| 9 | Kimi K3 Moonshot AI | 66 | 83 | 89 | 88 | 48 | - | 23 | 39 | - | 81 | 59 | 85 |
| 10 | Claude Sonnet 5 Anthropic | 65 | 73 | 91 | 87 | 38 | - | - | 29 | - | - | 54 | 80 |
How the Index Is Built
The index is the average of a model's scores on the benchmarks listed below, over the ones it has actually run: a benchmark it was never run on is not counted as a zero, it is simply absent. Every benchmark in the index counts equally. A benchmark that already reports a 0-100 percentage is used as it stands; one on a different scale is rescaled onto 0-100 first. A model needs results on at least two index benchmarks before it is ranked, so a single result does not give a model a standing of its own. Because scores are used at face value, a benchmark the whole field finds hard pulls every score on it down, and a model measured on harder benchmarks carries that into its average.
Benchmarks We Used
Recent Updates
Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.
Claude Opus 5.5
New model added to the AIMultiple Intelligence Index.
GPT-6 Sol
New model added to the AIMultiple Intelligence Index.
GPT-6 Luna
New model added to the AIMultiple Intelligence Index.
Grok 4.7
New model added to the AIMultiple Intelligence Index.
Explore LLM Use Cases, Analyses & Benchmarks
Audience Simulation: Can LLMs Predict Human Behavior?
In marketing, evaluating how accurately LLMs predict human behavior is crucial for assessing their effectiveness in anticipating audience needs and recognizing the risks of misalignment, ineffective communication, or unintended influence. Audience simulation with LLMs enables the modeling of virtual audiences, helping organizations anticipate reactions to content or products without relying on costly surveys or focus…
HALC-Bench: LLM Hallucination on Long-Context Retrieval Benchmark
HALC-Bench (LLM Hallucination on Long-Context Retrieval Benchmark) measures a large language model’s resistance to fabricating evidence for a metric that does not exist in the target document by using 3 haystacks placed at the beginning, middle, and end of the model’s context window, with 204 questions. claude-fable-5 answered all 204 traps correctly at every haystack…
AI Gateways for OpenAI: OpenRouter Alternatives
We benchmarked OpenRouter, SambaNova, TogetherAI, Groq, and AI/ML API across three indicators (first-token latency, total latency, and output-token count), with 300 tests using short prompts (approx. 18 tokens) and long prompts (approx. 203 tokens) for total latency. If you plan to use one of these AI gateways, you can: In this benchmark, we compared OpenRouter,…
LLM Observability Tools: Weights & Biases, Langsmith
LLM applications have expanded from single-turn chats into multi-step agents that use tools, query databases, and coordinate with other models, making their behavior harder to interpret. LLM observability provides continuous visibility into these complex workflows, helping organizations monitor quality, detect failures, troubleshoot issues, and manage performance and costs. W&B Weave is Weights & Biases‘ LLM…
Cloud LLM vs Local LLMs: Examples & Benefits
Cloud LLMs, powered by advanced models like GPT-5.5 and Claude Opus 4.7, offer scalability and accessibility. Conversely, Local LLMs, driven by open-source models such as Llama 4, DeepSeek V4, and Qwen3.6-Plus, ensure stronger privacy and customization. Explore what are cloud LLMs, strengths and weaknesses, most common case studies with real-life examples, and how they differ…
LLM Fine-Tuning Guide for Enterprises
Follow the links for the specific solutions to your LLM output challenges. If your LLM: The widespread adoption of large language models (LLMs) has improved our ability to process human language. However, their generic training often results in suboptimal performance for specific tasks. To overcome this limitation, fine-tuning methods are employed to tailor LLMs to…
10+ Large Language Model Examples
We have gathered open-source benchmarks to compare leading proprietary and open-source large language models. Choose your use case to find the right model. You can evaluate large language models by examining their benchmark performance and real-world latency (available by clicking each model’s name in the table), and by reviewing their pricing to assess overall efficiency…
LLM Pricing: Top 15+ Providers Compared
LLM pricing spans four orders of magnitude: the cheapest models launched under $0.03 per million tokens, while frontier reasoning tiers launched at up to $262.50. The chart below tracks launch prices: each point is the average price of the models one size class launched in a calendar quarter, blended 3 parts input to 1 part…
LLM Automation: Top 7 Tools & 8 Case Studies
LLM automation refers to shift to intelligent automation tools that leverage LLMs, including AI agents, fine-tuned LLMs and RAG models to automate and coordinate tasks. Explore what LLM automation is, its top real-life applications and major tools: Large language models in automation is a systematic approach that combines Natural Language Processing (NLP) with existing process…
LLM Orchestration: 22 Frameworks and Gateways
Optimizing LLM orchestration is key to improving performance while keeping resource use under control. To evaluate how different orchestration approaches perform in practice, we benchmarked: Discover selected LLM orchestration tools, including developer frameworks and enterprise gateways: LLM Orchestration involves managing and integrating multiple Large Language Models (LLMs) to perform complex tasks efficiently. It ensures smooth…