Premium
Services
Premium

AI VC Benchmark: 16 AI Agents on Client Identification

Ezgi Arslan, PhD.
Ezgi Arslan, PhD.
updated on Sep 10, 2026

Client identification is part of commercial due diligence. We tested 16 models on client-list research for three research and review sites, one of them AIMultiple, scoring their submissions against answer keys we compiled and the pages they cited.

Client identification benchmark results

Loading Chart

We did not set reasoning effort for any model. Each ran at its agent program’s default, and the scoring model was called at temperature 0.

  • Fable 5.1 scores 90.2, the highest average of the 16 models. Opus 5 follows at 89.1 and GLM 5.3 at 81.0.
  • GLM 5.3 scored between 79.6 and 83.3 on all three targets. Claude Sonnet 5 scored 90.8 on one and 0.0 on another, and Gemini 3.8 Flash scored 87.0 and 9.2. Averages of 48.0 and 60.7 describe neither model’s runs.

Target difficulty

Median scores are 64.3 on the first target, 59.2 on the second and 87.0 on the third. Ten of the 16 models score at least 84 on the third, whose answer key has 18 clients, so one correct client is worth 5.6 of the 100 coverage points that carry half the score.

Each model’s score is the mean of all three targets. Dropping an easy target after seeing the scores would bias the comparison. The next round needs harder target companies.

Cost and score comparison

  • Fable 5.1 averages 90.2 at $4.38 per task. Opus 5 averages 89.1 at $8.51, the highest cost of the 15 models with dollar figures. GLM 5.3 averages 81.0 at $1.40.
  • Grok 4.6 averages 80.5 at $0.42 per task, a twentieth of what Opus 5 costs for 8.6 fewer points.
  • GPT 6 Astra averages 74.2. Its three subscription tasks used 279,253 tokens, and Codex CLI records no dollar cost, so the cost chart omits it.

Why does client-list work test AI agents?

A client list requires a name and evidence of the commercial relationship. A plausible list can still cite pages that only compare products.

Public sources can confirm an individual client relationship. This benchmark does not test whether a list is complete.

A genuine client can be absent from the answer key. A submission earns credit for one when the cited page establishes the relationship, and loses coverage points for an unsupported extra company.

Get our team to automate one of your business processes with AI agents, free of charge.
Automate a process

Methodology

Model coverage

The score chart uses the 16 models selected for this benchmark campaign. Five superseded models stay in the historical data and are not charted. Scores run from 0 to 100 and bar labels round to whole points.

The agentic IT benchmark uses a different scale, so its scores are not comparable.

Tasks and submissions

The first target is AIMultiple, the publisher of this benchmark, and its answer key was compiled by an AIMultiple analyst. The second target is a software review site and the third an analyst research firm.

The three answer keys contain 45, 65, and 18 clients and were withheld from the tested models. The run covers 48 model-company assignments and 47 nonempty client lists; Inkling Small produced no client rows on the second company and scored zero there.

Each score uses one submission per company. The submissions span July through September, and some replace earlier attempts.

Reweighting the three targets across all 27 ordered draws with replacement puts the middle 95% of Fable 5.1’s means at 86.0-94.8 and Opus 5’s at 82.4-96.1, so Fable 5.1’s interval falls inside Opus 5’s. We reweighted only these three targets, and we did not measure run-to-run variation.

Costs and execution

Cost per task is the mean across a model’s three runs, including the run that produced no client rows. One task is one target company. Dollar costs cover 15 of the 16 models and elapsed-time data covers 13.

The dollar figures are not all metered the same way. Runs on opencode carry the OpenRouter charge for that session. Runs on Claude Code and Grok Build authenticate against a subscription, so their figures are list prices applied to the tokens the program reported. Codex CLI records no dollar cost.

Each model received a frozen prompt naming the company and requesting a CSV with company names, LinkedIn URLs, evidence URLs and evidence dates. Grok Build used native web search; the other agent programs reached the web through Bright Data, one of the web scraping tools we compare.

An agent program gives the model a shell and web access, so a score reflects both the model and the program it ran in. Provider list prices are in our LLM pricing comparison.

Scoring and evidence

Client coverage carries 50% of the score, LinkedIn accuracy 17%, evidence quality 17% and evidence recency 16%.

Scoring ran on September 10, 2026. It accepts evidence dated from September 10, 2025 to September 12, 2026, the 12 months before scoring plus two days. A cited URL returning 404 or 410 caps that submission at 40, and no submission hit the cap in this pass.

Code checks formatting, names, duplicates, dates and URL status. Claude Sonnet 5 reads an excerpt of each cited page and judges whether it confirms the relationship. Claude Sonnet 5 is also one of the 16 scored models, and Anthropic models hold the top two places. We did not test whether an Anthropic judge favours Anthropic models.

The judge prompt states whether the company is in the answer key, so the judge’s verdict is not independent of the key. We reran the scorer with one company dropped from a key, and on the same page and the same excerpt its verdict changed from client to mention.

The benchmark team recorded a decision in place of the judge’s verdict on 58 rows, 30 against the model and 28 for it. Fifty-four come from a September 9, 2026 review of the retained page text and four are editorial rulings recorded on September 10.

The evaluation accepts 140 rows outside the answer keys, 125 of them on the second target. Supported rows earn the same points as key entries.

Further readings

Don’t miss our benchmarks and data-driven insights. The button opens Google; selecting AIMultiple confirms that you wish to see AIMultiple more often in Google search results.
GoogleAdd as preferred source

FAQs

This test verified individual client relationships from public evidence. It did not assess whether a list is complete.

A comparison alone does not. A page must also establish a commercial relationship, such as a named customer or a subscriber to the research firm’s services.

Coverage values come from the answer key. The judge checks the page a model cited, so a submission earns credit for supported clients outside the key.

The test covers client-list research. It does not assess investment selection, valuation or the quality of a fund’s decisions.

Cite this benchmark

Pick the format that matches where you're publishing. Pasting the link version into your CMS preserves the backlink.

Ezgi Arslan, PhD. and Berk Kalelioğlu (2026) - "AI VC Benchmark: 16 AI Agents on Client Identification". Published online at AIMultiple.com. Retrieved September 10, 2026, from: https://aimultiple.com/ai-vc [Online Resource]

PhD., E. A., & Kalelioğlu, B. (2026, September 10). AI VC Benchmark: 16 AI Agents on Client Identification. AIMultiple. https://aimultiple.com/ai-vc

@misc{phd2026,
  author = {PhD., Ezgi Arslan, and Kalelioğlu, Berk},
  title  = {{AI VC Benchmark: 16 AI Agents on Client Identification}},
  year   = {2026},
  month  = sep,
  howpublished    = {\url{https://aimultiple.com/ai-vc}},
  note   = {AIMultiple. Retrieved September 10, 2026}
}
Download all data

Results and timestamps of 13 data points. Download the summary data shown in this article's charts and tables as a ZIP file containing one CSV file and a README.

Last updated: September 19, 2026
Download

Want the granular data behind it? Join Premium

Changelog

1 updates
  1. Replaced the deal sourcing and competitor mapping benchmark with a 14-agent client identification benchmark.

Ezgi Arslan, PhD.
Ezgi Arslan, PhD.
Industry Analyst
Ezgi holds a PhD in Business Administration with a specialization in finance and serves as an Industry Analyst at AIMultiple. She drives research and insights at the intersection of technology and business, with expertise spanning sustainability, survey and sentiment analysis, AI agent applications in finance, answer engine optimization, firewall management, and procurement technologies.
View Full Profile
Technically reviewed by
Berk Kalelioğlu
Berk Kalelioğlu
AI Researcher
Berk is an AI Researcher at AIMultiple's benchmark team, focusing on agentic AI, machine learning, and large and small language models (LLMs & SLMs).
View Full Profile

Be the first to comment

Your email address will not be published. All fields are required. Comments are left in their original language.

0/450