Enterprises use LLMs every day for their regular tasks. To find the most cost-efficient LLMs, we designed AIM Enterprise, an agentic enterprise benchmark, where we used 69 real enterprise tasks across strategy, marketing, HR, sales, and operations.
Benchmark results
- Opus 5 scored higher on 39 tasks and Astra on 30. The top four, with Claude Fable 5.1 and Grok 4.6, span 2.9 points.
- Astra led on 23 tasks, Opus 5 on 22, and Fable 5.1 on 16.
- Claude Fable 5.1 beat Claude Fable 5 on 63 of 69 tasks and averaged 7.2 points more. GPT 6 Astra averaged 8.8 points more than GPT 5.6 Sol and beat it on 61.
- Claude Sonnet 5 costs $1.23 a task. Gemini 3.8 Flash scored 1.5 points higher at $0.61, and Grok 4.6 scored 13 points higher at $1.09.
Cost and score
- GPT 6 Astra and Claude Opus 5 cost about the same, $1.76 and $1.73 a task, and score 1.6 points apart.
- GPT 5.6 Luna scored 45.2 at $0.016 a task, 18.7 points below Claude Opus 5 at under one percent of its cost.
- GPT 5.6 Sol scores 53.5 at $0.167 a task. Grok 4.6 scores 7.5 points higher and costs 6.5 times as much.
- Claude Fable 5.1 is the most expensive model on the chart at $3.19 a task and scores 1.8 points below Claude Opus 5, which costs $1.73.
- Across the 16 models on the chart, cost per task and score correlate at 0.68. Median time per task and score correlate at 0.04. The fastest model, Inkling Small at 66 seconds a task, ranks last.
Judge disagreement
Two judge models scored every file that passed the checks. On 32.7% of the individual scores the two differed by more than a quarter of the scale, and nobody has reviewed those rows.
Both judges put the same four setups in the top four and the same two last; 12 of the 17 setups rank differently under each judge.
Each judge put its own company’s model first. GPT 5.6 Sol ranked GPT 6 Astra first and Claude Opus 5 second; Opus 5 ranked itself first and Astra fourth. The published order combines both rankings.
AIM Marketing benchmark
The same method runs on marketing work: finding offerings a competitor publishes and AIMultiple does not, building a best-fit account list, and producing a personalized sales deck. Each task scores 0 to 100 and the overall score is the mean of the three. A website reputation audit runs alongside them and is reported on its own axes, because it has no fixed maximum score.
Full results: the agentic marketing benchmark.
AIM IT benchmark
Twelve models each ran a benchmark-design task twice, inventing a benchmark, building it, and running four models through it. None of the 24 attempts passed every criterion, and six of the rubric’s checks were passed by none of them. Claude Opus 5 led on text-to-SQL at 78.2 and Kimi K3 on tool calling at 74.3.
Full results: can LLMs design a benchmark.
AIM VC benchmark
Thirteen charted models were asked to name a company’s clients with dated evidence, across three target companies. Claude Opus 5 scored 89.1 of 100 and Claude Fable 5 85.0, and those two were the only ones that stayed above 76 on all three targets as scored. Most of the others scored lowest on the target whose clients appear in podcast advertising rather than on indexed pages.
Full results: the client-list benchmark.
Agentic enterprise platforms
Companies can also use LLMs through business software. The platforms below connect models from providers such as OpenAI, Anthropic, and Google to CRM records, workflows, and prebuilt agents.
Creatio
Creatio is a CRM and workflow platform with prebuilt AI agents.
Agents by department: The agents available depend on the apps a company has installed. Examples:
- Sales Agent: enriches account data, fills in opportunity fields after Zoom meetings, and prepares for new meetings.
- Knowledge Agent: drafts knowledge base articles from resolved cases.
- Segmentation Agent: turns plain-language audience descriptions into marketing segments.
A separate model per agent, including self-hosted ones: Each agent or sub-agent can use its own model. Supported options are OpenAI, Azure OpenAI, and LiteLLM-supported providers, including models on the company’s own infrastructure. Agent scenarios require at least a 40,000-token context window and function calling.
Salesforce Agentforce
Agentforce is the agent layer of Salesforce’s CRM.
Models: Agents use Salesforce-enabled models from partners, including Anthropic, Google, and OpenAI. Companies can also connect their own LLM account, and an agent’s default model can be changed.
Einstein Trust Layer: Model calls pass through a layer that Salesforce says includes toxicity detection, zero data-retention agreements, and prompt-injection defense.
Handoff to a person: “Transfer to human agent” is one of the actions an agent can choose. This lets a conversation move to a person when a process cannot tolerate errors.
Microsoft Copilot Studio
Copilot Studio is Microsoft’s tool for building agents for Microsoft 365 and Power Platform.
Models: OpenAI models are the default. Creators can also add external models from Anthropic, xAI, or Mistral.
Production labels: Microsoft marks some models as experimental or preview and advises against using them in production. Published agents that use these models are still billed at standard rates.
Fallback: If an admin turns off Anthropic models, agents built on them switch to the default OpenAI model automatically.
Methodology
The 69 tasks were written by AIMultiple’s founder against real company decisions and mapped to APQC Process Classification Framework IDs.1
They cover strategy (16 tasks), marketing (15), HR (7), sales (6), operations (5), IT and finance (4 each), and seven smaller areas. Each task names the file it wants: a results.csv with a fixed column list, typically 10 rows and 8 columns, with an explicit rule for every integer field.
How the models ran
Sixteen models ran in 17 setups, a setup being one model under one agent program. Eight ran on opencode 1.15.13 through OpenRouter. The other nine used Claude Code, Codex, or Grok Build, their vendors’ own agent programs. All of them billed against subscriptions.
Fable 5.1 ran on Claude Code 2.1.258 while the other three Claude setups ran on 2.1.220; Astra ran on Codex 0.153.4 while the three GPT 5.6 setups ran on 0.146.0. Gemini 3.8 Flash ran on the same opencode build as the rest.
Every setup got the same frozen prompt, live web access through a scraping API and a one-hour limit. Prompts were never tuned per model, and the scoring rubrics never reached the machine that ran the tasks.
Nobody chose a reasoning level. Every setup ran at its agent program’s default, and the defaults are not one level:
Claude Code defaults to high on every model it ran. Codex sends each model’s catalog default. That is low for GPT 5.6 Sol and medium for Luna, Terra, and GPT 6 Astra. Grok Build ran Grok 4.6 at high. opencode leaves the level to the provider, so its eight setups ran at each model’s OpenRouter default.
A higher reasoning default did not come with a higher score. On the 16-model chart, the three models whose default is above high rank 9th, 11th and 14th: GLM 5.3 at max, Qwen 3.8 Max at xhigh and Kimi K3 at max. Claude Sonnet 5 ran under both a vendor CLI and opencode and scored 47.9 and 47.5.
Every cost figure here is recomputed from recorded token use at OpenRouter list prices frozen on August 27, 2026. Reasoning tokens are priced as output tokens, which is how OpenRouter bills them. GPT 6 Astra and Gemini 3.8 Flash joined later and use OpenRouter rates read on September 5, 2026; Claude Fable 5.1 uses Anthropic’s published rates of the same date, because OpenRouter does not list it.
The Claude Code costs also apply three published Anthropic rules that a flat price list lacks. The one-hour prompt cache costs twice the input rate, where a five-minute cache costs 1.25 times the input rate. Web search costs $10 per thousand requests. The Haiku 4.5 subagent that Claude Code runs alongside the main model is included.
The agent programs’ own cost figures are not comparable. Every run was on a subscription, so no setup has a per-run bill. Where a program reports a cost, it prices its own tokens against its own list. opencode uses a bundled price list. Grok Build uses an xAI list at about a third of the public rate. Codex reports nothing. Claude Code priced Sonnet 5 at $3 and $15 per million tokens, Claude Sonnet 4.6’s rate, where Sonnet 5 is $2 and $10.
The runs produced 1,140 files out of 1,173 setup-task pairs. A missing file scores zero and is not rerun, unless the run hit the one-hour limit or ended on an error from the model’s provider; each exception allows one retry.
Five runs stalled and four delivered on the retry. Eighteen runs ended on a provider error, a 502 or 504 from the API or a corrupted response. One had already written its file; the other 17 were rerun, and 16 delivered. One run hit both the time limit and a provider error.
Eight Qwen 3.8 Max runs ended when opencode’s loop guard stopped the model from repeating the same tool call a third time. No other setup triggered the guard. Those runs were not rerun, and four of them produced no file. Every file was scored once, so no result came from picking the better of two attempts.
How the checks and judges work
Deterministic checks run first, and no judge sees a file that fails them: column set and row count, RFC 4180 parsing,2
integer formatting, no empty cells, no duplicate rows, and sort order where the task requires one.
A file that fails scores zero rather than being dropped. MiniMax M3 failed on 8 of the 64 files it delivered, six of them because the CSV would not parse, and GPT 5.6 Luna and GPT 5.6 Terra failed on one each. Removing every task where any setup scored zero leaves 38 tasks and the same leader.
The 1,130 files that passed went to two judges: GPT 5.6 Sol through the Codex CLI and Claude Opus 5 through the Claude Code CLI, both at high reasoning effort with live web access.
Each column of the file goes to its own subagent, which sees that column’s answers from every model and nothing else, under anonymous row IDs shuffled separately for each judge. The judge ranks those answers against each other and cannot call any two equal. The two rankings are then combined by adding up each answer’s position under each judge. That produced 93,490 scores.
Measured with model names visible, Sol ranked GPT answers 8.4 percentile points above where Opus ranked them, and Opus ranked Anthropic answers 4.9 percentile points above where Sol ranked them. In the anonymized run behind this leaderboard, Sol ranked GPT 6 Astra first and Opus ranked Claude Opus 5 first.
How scoring works
A model’s score is the sum of its column averages, rescaled so the highest possible total is 100. That ceiling depends on how many models passed the checks, which is why these scores compare inside this benchmark and nowhere else.
To test how much the ranking depends on the task selection, we rescored the leaderboard 4,000 times, each time on a random re-selection of the 69 tasks. That measures task sensitivity only. It does not measure how much a rerun of the same task would move a score, because no task was scored twice.
The hardest quarter is 17 tasks, the ones with the lowest average score across all 17 setups. Five tasks tie for the last four places, and Opus 5 has the highest average under all five possible selections. That rule was fixed before any model’s standing on those tasks was calculated.
Ten of the 16 models also appear in the benchmarks above. Their scores are not comparable across benchmarks, because each sets its own scale.
Every setup here is scored on its position relative to the models it ran against, so dropping one setup changes every other score. For the same reason these scores do not compare with the previous run of this benchmark, which scored a different field. Only the order of the setups can be compared across runs.
Gemini 3.7 Flash ran on this field and was dropped when Gemini 3.8 Flash completed, under the rule that a model leaves once its successor has run. The same rule retired Grok 4.5 and GLM 5.2.
FAQs
Not on this evidence. The scale is relative, so even the top score of 63.9 means only that a model’s answers ranked above the others, and the rows where the two scoring models substantially disagree are unreviewed. Vendors building this kind of agent are listed in our enterprise AI companies breakdown, and AIMultiple automates processes like these.
Seven of them score between 45.2 and 49.4, closer than 69 tasks can separate. The top four are also within 2.9 points of each other. The benchmark separates strong models from weak ones; resampling the tasks reorders both of those groups.
One model ran under both programs on this run, and it scored 47.9 under Claude Code and 47.5 under opencode. One comparison does not give a general answer. The charts show the Claude Code result. The opencode setup read 4.5 times as much cached context and produced 1.2 times as many output tokens, including reasoning. It cost $1.59 per task, compared to $1.23 for Claude Code, whose figure already includes a subagent that opencode does not run.
Cite this benchmark
Pick the format that matches where you're publishing. Pasting the link version into your CMS preserves the backlink.
@misc{kalelioglu2026,
author = {Kalelioğlu, Berk},
title = {{AIM Enterprise: Agentic Enterprise Benchmark}},
year = {2026},
month = sep,
howpublished = {\url{https://aimultiple.com/agentic-enterprise}},
note = {AIMultiple. Retrieved September 17, 2026}
}Results and timestamps of 33 data points. Download the summary data shown in this article's charts and tables as a ZIP file containing 2 CSV files and a README.
Want the granular data behind it? Join Premium
Changelog
1 updatesAdded a table of reasoning effort by model and agent program to the methodology section.
Be the first to comment
Your email address will not be published. All fields are required. Comments are left in their original language.