Services
Contact Us

AIM Enterprise: Agentic Enterprise Benchmark

Berk Kalelioğlu
Berk Kalelioğlu
updated on Aug 24, 2026

Enterprises use LLMs every day for their regular tasks. To find the most cost-efficient LLMs, we designed AIM Enterprise, an agentic enterprise benchmark, where we used 69 real enterprise tasks across strategy, marketing, HR, sales, and operations.

Benchmark results

Loading Chart
  • Claude Opus 5 (66.4) and GPT 5.6 Sol (62.4) finished more than 7 points ahead of third place, and held the top two spots in all 4,000 rescorings on random re-selections of the tasks.
  • Opus 5 won 43 of the 69 tasks and Sol won 13. No other model won more than four.
  • Opus 5 also had the highest average on the 17 hardest tasks, 66.2 against 62.3 for Sol.
  • The next six models scored between 51.4 and 55.2. That is closer than 69 tasks can separate, so their order moves with the task selection.

Cost and score

  • GPT 5.6 Luna scored 55.2 at $0.040 per task. That is 7 points below the best model on the chart at a seventeenth of its cost.
  • Across the 11 priced models, cost per task and score correlate at 0.71.
  • Across all 16 setups the correlation between time per task and score is 0.55. The four fastest were four of the five lowest scorers, and Inkling Small finished a task in 49 seconds to come second from bottom.
  • Kimi K3 at $0.652 and Qwen 3.8 Max at $0.628 cost about what GPT 5.6 Sol does at $0.670, and score 8 and 11 points lower.

Judge disagreement

Two judge models scored every file, and they often disagreed. On about a third of the individual scores the two landed more than a quarter of the scale apart, and nobody has reviewed those rows.

They agreed on the ends and not the middle. Both put Opus 5 and GPT 5.6 Sol above everyone else, but 11 of the 16 setups land in a different position depending on which scoring model you follow.

The two also disagreed on which leader comes first, each placing their company’s model on top. The published order combines both rankings.

AIM Marketing benchmark

The same method runs on marketing work: finding offerings a competitor publishes and AIMultiple does not, building a best-fit account list, and producing a personalized sales deck. Each task scores 0 to 100 and the overall score is the mean of the three. A website reputation audit runs alongside them and is reported on its own axes, because it has no fixed maximum score.

Full results: the agentic marketing benchmark.

Get our team to automate one of your business processes with AI agents, free of charge.
Automate a process

AIM IT benchmark

Twelve models each ran a benchmark-design task twice, inventing a benchmark, building it, and running four models through it. None of the 24 attempts passed every criterion, and six of the rubric’s checks were passed by none of them. Claude Opus 5 led on text-to-SQL at 78.2 and Kimi K3 on tool calling at 74.3.

Full results: can LLMs design a benchmark.

AIM VC benchmark

Thirteen charted models were asked to name a company’s clients with dated evidence, across three target companies. Claude Opus 5 scored 89.1 of 100 and Claude Fable 5 85.0, and those two were the only ones that stayed above 76 on all three targets as scored. Most of the others scored lowest on the target whose clients appear in podcast advertising rather than on indexed pages.

Full results: the client-list benchmark.

Don’t miss our benchmarks and data-driven insights. The button opens Google; selecting AIMultiple confirms that you wish to see AIMultiple more often in Google search results.
GoogleAdd as preferred source

Methodology

The 69 tasks were written by AIMultiple’s founder against real company decisions and mapped to APQC Process Classification Framework IDs.1

They cover strategy (16 tasks), marketing (15), HR (7), sales (6), operations (5), IT and finance (4 each), and seven smaller areas. Each task names the file it wants: a results.csv with a fixed column list, typically 10 rows and 8 columns, with an explicit rule for every integer field.

How the models ran

Thirteen models ran in 16 setups, a setup being one model under one agent program. Eleven ran on opencode 1.15.13 through OpenRouter. The rest ran on the agent programs their own vendors ship, Claude Code and Codex, billed against subscriptions.

Every setup got the same frozen prompt, live web access through a scraping API and a two-hour limit. Prompts were never tuned per model, and the scoring rubrics never reached the machine that ran the tasks.

Nobody chose a reasoning level. Every setup ran at its agent program’s default, and the defaults are not one level:

Claude Code 2.1.220 ships high for both models it ran and Codex recorded high on every run. opencode chooses nothing and leaves the level to the provider, so those eight sit at each model’s own default.

A higher reasoning default did not track a higher score. The two models that default above high, Kimi K3 at max and Qwen 3.8 Max at xhigh, finished 4th and 8th of 13. The three models that ran under both a vendor CLI and opencode land within 0.83 points of themselves.

Every cost figure here is recomputed from the tokens a run actually moved, at OpenRouter list prices read on August 19, 2026, with cached input billed at the cached rate.

The agent programs’ own billing would not compare. opencode bills at its own bundled price list, and the Claude Code and Codex runs bill against subscriptions that record no charge per run. The two Claude Code setups kept no usage record and cannot be priced at all.

That produced 1,104 files, one per setup per task. Six opencode setups missed 25 runs on the first pass because the agent’s file search walked paths its own sandbox then refused to open, which stalled the run. Rerunning those 25 in an isolated directory recovered all of them, so delivery is reported as both a first-pass and a final rate. Every task ran once and was scored once. A model that produced no file was run again, but no file was ever scored twice, so no result came from picking the better of two attempts.

A seventieth task, an inquiry-handling playbook, is left out of every figure here. Its input file was missing at run time and only one setup ever produced a file for it. Counting every restart, 83 of the 759 runs with usage records took more than one try.

How the checks and judges work

Deterministic checks run first, and no judge sees a file that fails them: column set and row count, RFC 4180 parsing,2

integer formatting, no empty cells, no duplicate rows, and sort order where the task requires one.

A file that fails scores zero rather than being dropped. DeepSeek V4 Flash failed on 6 of its 69 files, Inkling Small on 2, MiniMax M3 and GPT 5.6 Terra on one each. Removing every task where any setup failed leaves the two leaders in place.

The 1,094 files that passed went to two judges: GPT 5.6 Sol through the Codex CLI and Claude Opus 5 through the Claude Code CLI, both at high reasoning effort with live web access.

Each column of the file goes to its own subagent, which sees that column’s answers from every model and nothing else, under anonymous row ids shuffled separately for each judge. The judge ranks those answers against each other and cannot call any two equal. The two rankings are then combined by adding up each answer’s position under each judge. That produced 90,530 scores.

Measured with model names visible, Sol ranked GPT answers 8.4 percentiles above where Opus ranked them, and Opus ranked Anthropic answers 4.9 percentiles above where Sol ranked them. Anonymization does not remove the effect. The runs behind this leaderboard were anonymized, and each judge still put its own company’s setup first.

How scoring works

A model’s score is the sum of its column averages, rescaled so the highest possible total is 100. That ceiling depends on how many models passed the checks, which is why these scores compare inside this benchmark and nowhere else.

To test how much the ranking depends on the task selection, we rescored the leaderboard 4,000 times, each time on a random re-selection of the 69 tasks. That measures task sensitivity only. It does not measure how much a rerun of the same task would move a score, because no task was scored twice.

The hardest quarter is the 17 tasks with the lowest average score across all 16 setups. That rule was fixed before any model’s standing on those tasks was calculated.

Twelve of the 13 models also appear in the benchmarks above. Their scores do not carry across, because each benchmark sets its own scale. 

Every score here is a position relative to the models it ran against, so dropping one setup would rescore all the rest instead of removing a bar. Qwen 3.8 Max runs this benchmark only and stays in the chart for that reason.

FAQs

Not on this evidence. The scale is relative, so even the top score of 66.4 says only that a model’s answers ranked above the others, and the rows where the two scoring models substantially disagree are unreviewed. Vendors building this kind of agent are listed in our enterprise AI companies breakdown, and AIMultiple automates processes like these.

Because six of them are genuinely close, and 69 tasks cannot separate differences under about 1.5 points. The benchmark separates strong models from weak ones. Task resampling reorders the middle six.

Not for output quality, in the three cases we could test. Three models each ran under two agent programs and no pair differed by more than 0.83 points, even though the vendor CLIs reasoned at high and the opencode runs took the provider default. The charts draw the vendor’s own program for those three. The difference showed up in operations instead: only the opencode setups lost runs to a sandbox conflict, and only Claude Code left no usage record.

Cite this benchmark

Pick the format that matches where you're publishing. Pasting the link version into your CMS preserves the backlink.

Berk Kalelioğlu (2026) - "AIM Enterprise: Agentic Enterprise Benchmark". Published online at AIMultiple.com. Retrieved August 24, 2026, from: https://aimultiple.com/agentic-enterprise [Online Resource]

Kalelioğlu, B. (2026, August 24). AIM Enterprise: Agentic Enterprise Benchmark. AIMultiple. https://aimultiple.com/agentic-enterprise

@misc{kalelioglu2026,
  author = {Kalelioğlu, Berk},
  title  = {{AIM Enterprise: Agentic Enterprise Benchmark}},
  year   = {2026},
  month  = aug,
  howpublished    = {\url{https://aimultiple.com/agentic-enterprise}},
  note   = {AIMultiple. Retrieved August 24, 2026}
}
Berk Kalelioğlu
Berk Kalelioğlu
AI Researcher
Berk is an AI Researcher at AIMultiple's benchmark team, focusing on agentic AI, machine learning, and large and small language models (LLMs & SLMs).
View Full Profile

Be the first to comment

Your email address will not be published. All fields are required. Comments are left in their original language.

0/450