Commercial due diligence starts by working out who pays the target. We tested 14 models on that question, each asked to build a client list for three real companies and scored against an answer key our analyst compiled by hand.
Client identification benchmark results
Reasoning effort was not set for any model. Each ran at its agent program’s default, and the scoring model was called at temperature 0.
- Claude Opus 5 averaged 89.1 and Claude Fable 5 85.0, and as scored neither fell below 83.3 and 76.9 on any single company. The other twelve each fell under 60 somewhere, and eight fell under 30.
- The second target scored lowest, with a median of 29.4 and six models at or below 4.6. Its client evidence sits in buyer’s guides and reviews rather than case studies. Those same six models score a median 54.9 on the third target.
- Model means span 77.0 points against 39.1 for target means, but the order still reshuffles between targets. MiniMax M3 ranks 12th of 14 on the first target and 4th on the second.
- Ninety-three rows are still with human reviewers, most of them because the page a model cited does not confirm a client the answer key names. Counting all 93 as wrong drops Opus 5 to 74.6, still first. The 14 models lose 9.9 points on average and Opus 5 loses the most, 14.6.
The third target needs replacing
Seven of 14 models score 84 or above on the third target and Opus 5 scores 100. Its answer key has 18 clients, so one correct client is worth 5.56 of the 100 coverage points and a model that finds most of them lands near the top.
Those scores are also the ones most likely to change. Sixty-seven of the 93 disputed rows come from that target, and counting all of them as wrong leaves nobody above 84 there and Opus 5 at 66.7. An 18-client key is too small either way, which is why the company should be replaced.
We kept it in the average. Dropping it after seeing the results would bias the average. The three targets still separate the models, but the next round needs a harder third company.
Cost and score comparison
The top score costs the most. Opus 5 runs $8.52 per task, seven times the $1.18 median across the 14 models, and the next most expensive, Claude Fable 5, runs $4.97.
Across the 14 models the correlation between cost per task and score is 0.89. Sol scored 71.2 at $3.25 per task, 18 points below Opus 5 for 62% less money. Kimi K3 scored 56.5 at $2.98.
Two models tie for third. Qwen 3.8 Max scored 71.2 against Sol’s 71.2 at $2.16 per task, a third less than Sol paid for the same result.
Price does not always win. MiniMax M3 at $0.18 per task (52.7) and Grok 4.5 at $0.58 (49.4) both beat Claude Sonnet 5 at $1.71 (48.9) and Gemini 3.6 Flash at $1.77 (48.7).
Slower runs score higher. Inkling Small was the fastest at 173 seconds and the lowest scorer at 12.1, while Opus 5 took 885 seconds. The correlation between time and score is 0.78, and 0.71 with the best and worst model removed.
Task selection
Building a client list is the part of commercial due diligence that can be checked against public evidence. Three properties make it a benchmark task.
Funds and advisers pay analysts to build these lists today.
Answer-key entries are checkable against public pages, and a relationship with no public trace is outside what the benchmark can score either way.
A model that cannot find a client can name a plausible one instead. Both the key and the judge score those rows negative.
Methodology
Thirteen of the 14 charted models also ran the benchmark-design benchmark and the enterprise-task benchmark, so the same models can be compared across the three. Scores do not compare across the three, because each benchmark uses its own scale.
Claude Fable 5 ran this benchmark only and is charted anyway. One further model, a preview build of Gemini 3.1 Pro, ran the task and is in the data file but not the chart.
Cost is the OpenRouter charge for each scored run, metered separately by a proxy in front of the API so runs going on at the same time never mix. Inkling Small’s one non-delivering run is included in its cost. Provider list prices are in our LLM pricing comparison.
The three targets are a technology research and comparison site, a software review site and an analyst research firm.
In that order, every row a model delivered was checked against answer keys of 45, 65 and 18 clients, each compiled by hand by one of our analysts and withheld from the models. Every model-company score comes from a single run, so run-to-run variation is unmeasured and gaps of a few points do not mean much. Inkling Small’s one non-delivering run scores zero.
Eight of the 42 scored runs had a second attempt on file, and the attempt with more rows was kept. Two of those pairs were tied and two of the alternatives were truncated, leaving four where both attempts were complete. Kimi K3 and GPT 5.6 Sol are affected on all three targets, Claude Fable 5 and Gemini 3.6 Flash on the third target. Claude Opus 5 ran once everywhere.
Resolving all 93 disputed rows against the models leaves the top five and bottom four in place; only positions six to ten swap, so a 5.6-point gap in the middle of the table does not establish an order.
The cost figures quoted above are the mean for each model across its three runs. Time figures are medians. Correlations are Spearman rank correlations, which compare the order of two lists rather than the raw numbers.
Each model received one prompt per company naming the target and a required CSV schema (company, LinkedIn URL, evidence URL, evidence date). Every model had live web access through Bright Data’s MCP server, one of the web scraping tools we compare.
Runs used the opencode agent program at version 1.15.13 on a single server, with a three-hour limit per run.
Scoring is 50% client coverage against the answer key, 17% LinkedIn accuracy, 17% evidence quality and 16% evidence recency, with a 12-month recency window anchored at 2026-07-30. A row citing a URL that returns 404 or 410 caps the whole file at 40. No file triggered the cap. A URL that fails to load for any other reason goes to review instead of counting as fabrication.
Format checks, answer-key matching, duplicate detection, date parsing and URL liveness are deterministic. Claude Sonnet 5 reads the fetched page and judges whether a company outside the answer key is a genuine client, and whether the cited page confirms the relationship it is offered as evidence for.
The judge accepted 82 out-of-key rows across the 14 charted models, 69 of them on the second target. An accepted out-of-key row earns the same points as a keyed client, so the key sets how much one client is worth rather than the full list of clients that count.
Further readings
- AIM Agentic Marketing Benchmark
- AI-Based Stock Trading: Which Gen AI Tool Is Better
- Agentic AI Finance Benchmark
FAQs
No. The tasks map to work junior analysts do: sourcing, competitor mapping, comps. Investment decisions, partner judgment, and founder relationships sit outside their scope. AI compresses research time; humans make the decision.
Public companies have audited financials, which makes valuation a lookup exercise. Private early-stage startups force the agent to estimate, state a basis, and label uncertainty. That is where fabrication pressure is highest and where real due diligence effort is put in.
Do these benchmarks show AI replacing venture capital analysts?
No. The tasks map to work junior analysts do: sourcing, competitor mapping, comps. Investment decisions, partner judgment, and founder relationships sit outside their scope. AI compresses research time; humans make the decision.
Why do the tasks target private companies?
Public companies have audited financials, which makes valuation a lookup exercise. Private early-stage startups force the agent to estimate, state a basis, and label uncertainty. That is where fabrication pressure is highest and where real due diligence effort is put in.
Cite this benchmark
Pick the format that matches where you're publishing. Pasting the link version into your CMS preserves the backlink.
@misc{phd2026,
author = {PhD., Ezgi Arslan, and Kalelioğlu, Berk},
title = {{AI VC Benchmark: 14 AI Agents on Client Identification}},
year = {2026},
month = aug,
howpublished = {\url{https://aimultiple.com/ai-vc}},
note = {AIMultiple. Retrieved August 27, 2026}
}Results and timestamps of 0 data points. Download the data used in this article as a ZIP file containing 0 CSV files.
Be the first to comment
Your email address will not be published. All fields are required. Comments are left in their original language.