LLM Benchmarks
One transparent Intelligence Index combining public benchmarks with AIMultiple's own agentic, RAG and enterprise-reasoning evaluations.
Leaderboard
The highest-scoring models across all benchmarks.
# | Model | Index | LegalBench | FinanceReasoning | Text-to-SQL | Agentic RAG | ARC-AGI-2 | Agentic LLM Benchmark | FrontierMath | Swe-Bench |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 Anthropic | 92 | 89 | 90 | 90 | 98 | - | 69 | - | - |
| 2 | Kimi K3 Moonshot AI | 88 | 86 | 88 | 77 | 93 | - | 73 | - | - |
| 3 | GPT-5.6 Sol OpenAI | 87 | 87 | 90 | 74 | 87 | 93 | 62 | - | - |
| 4 | Grok 4.5 X | 86 | 86 | 88 | 79 | 83 | - | 73 | - | - |
| 5 | GPT-5.5 OpenAI | 83 | 87 | - | - | - | 85 | 59 | - | - |
| 6 | GPT-5.6 Terra OpenAI | 80 | 85 | 87 | 71 | 91 | 84 | 61 | - | - |
| 7 | GPT-5.6 Sol Pro OpenAI | 80 | - | 91 | 79 | 96 | - | 54 | - | - |
| 8 | Claude Opus 4.6 Anthropic | 78 | - | 88 | 68 | 80 | 69 | 72 | - | - |
| 9 | Gemini 3 Pro Preview Google | 77 | 87 | 86 | 60 | 89 | 31 | - | - | - |
| 10 | Gemini 3.1 Pro Preview Google | 77 | 87 | 87 | 65 | 89 | 77 | 46 | 27 | - |
Page 1 of 7 | ||||||||||
Charts
Following "Best Performing" models
How The Index is Built
The Intelligence Index normalizes each benchmark to a 0-100 scale and averages a model's relative standing across the benchmarks it was evaluated on. Each benchmark below feeds that score.
Recent Updates
Latest changes to the Intelligence Index, model coverage and AIMultiple benchmark methodology.
Kimi K3
New model added to the AIMultiple Intelligence Index.
GPT-5.6 Sol
New model added to the AIMultiple Intelligence Index.
GPT-5.6 Terra
New model added to the AIMultiple Intelligence Index.
GPT-5.6 Sol Pro
New model added to the AIMultiple Intelligence Index.
Explore LLM Use Cases, Analyses & Benchmarks
AI Gateways for OpenAI: OpenRouter Alternatives
We benchmarked OpenRouter, SambaNova, TogetherAI, Groq, and AI/ML API across three indicators (first-token latency, total latency, and output-token count), with 300 tests using short prompts (approx. 18 tokens) and long prompts (approx. 203 tokens) for total latency. If you plan to use one of these AI gateways, you can: In this benchmark, we compared OpenRouter,…
Large Language Models in Cybersecurity
We evaluated 7 large language models across 9 cybersecurity domains using SecBench, a large-scale and multi-format benchmark for security tasks. We tested each model on 44,823 multiple-choice questions (MCQs) and 3,087 short-answer questions (SAQs), covering data security, identity & access management, network security, vulnerability management, and cloud security. MCQs (Multiple-Choice Questions) benchmarking: SAQs (Short Answer…
ChatGPT for Customer Service: Top 10 Use Cases
ChatGPT has moved from novelty to infrastructure in customer service. Companies are using it to cut response times, handle volume their teams can’t absorb, and reduce the cost of routine interactions. But results vary sharply depending on how it’s implemented. OpenAI launched GPT-5.2, a materially more capable model that is better at instruction-following, reasoning across…
LLM Quantization: BF16 vs FP8 vs INT4
We benchmarked Qwen3-32B at 4 precision levels (BF16, FP8, GPTQ-Int8, GPTQ-Int4) on a single NVIDIA H100 80GB GPU. Each configuration was evaluated on 2 benchmarks (~12.2K questions) covering knowledge and code generation, plus 2,000+ inference runs to measure throughput. Int4 is 2.7x faster than BF16 while losing less than 2 points on MMLU-Pro, but code…