We are introducing AIM Agentic Marketing Benchmark which measures agent performance in competitive gap analysis and ABM target list preparation.
We tested the performance of 11 models and measured end-to-end execution performance:
Agentic marketing benchmark results
The task scores are normalized to a 0–100 scale.
- For competitive analysis, the score equals: 100 × MAX(0, verified gaps found − incorrect gaps) / 10
- For account-list preparation, the score is the percentage of 71 available rubric points earned by the model.
The overall score is the arithmetic mean of the two task scores, giving each task equal weight despite their different rubric sizes. Invalid or missing submissions count as zero in the overall calculation, although the table marks them separately from valid submissions that happen to score zero.
Read the methodology for more information on task evaluation.
Fable 5 ranked first in both tasks
Fable 5 achieved the highest overall score at 75.15, with 70 in competitive analysis and 80.3 in account-list preparation. Its lead came from producing the strongest valid result on both workflows and avoiding the zero that follows an invalid or missing submission.
The aggregate scores do not reveal which underlying capability created this advantage. Format compliance, source selection, search coverage, and repeated rule application could all contribute; separating them would require a criterion-level comparison of each model’s files.
GLM 5.2 ranked second with 68.05 and also placed second on both tasks. Fable 5 led by 10 points in competitive analysis and by 4.2 points in account-list preparation.
Competitive analysis returned lower scores
When the invalid submission is included as zero, the mean competitive-analysis score is 40.9. Fable 5 led with 70, while valid scores ranged from 30 to 70.
These figures need to be interpreted alongside the scoring formula. Because the task uses ten expected gaps, each additional net correct finding changes the score by ten points. A model that finds six valid gaps and no false ones scores 60; a model that finds seven valid gaps and adds one incorrect claim receives the same result.
The rubric also describes this task as a weak discriminator among frontier models. Its main value lies in separating weaker systems that miss several gaps, add false absence claims, or fail the required output format. Small score differences near the top should therefore receive less weight than larger failures or invalid submissions.
Account-list preparation provided more score resolution
The account-list task uses a 71-point rubric, which produces more granular results. Valid scores ranged from 32.4 to 80.3, and Fable 5 and GLM 5.2 exceeded 70.
Sol, Terra, Sonnet 5, and Opus 4.8 formed a close middle group, scoring between 66.2 and 69. Gemini 3.1 Pro returned the lowest valid submission at 32.4. Kimi K3 did not produce a scored result, while MiniMax M3 failed a mandatory structural condition.
The average was 51.5 after invalid and missing runs were converted to zero. Because the workflow required a coverage map and 100 qualified accounts, the score reflected performance across hundreds of individual claims and classifications rather than a short list of findings.
Performance varied by workflow
Kimi K3 scored 60 in competitive analysis and did not return a valid account list. Sonnet 5 showed the reverse pattern: its competitive-analysis output was invalid, while its account-list submission scored 67.6.
These cases show why the task-level results matter. Competitive analysis emphasizes semantic comparison and absence verification; account-list preparation places heavier demands on sustained research, entity resolution, source management, and file construction.
For marketing teams, the practical implication is clear: choose models based on the workflow being automated, then test them against the required output format and evidence standard.
The real-life tasks used in the benchmark
These agentic marketing tasks include strategy and content coverage, growth, and sales research. Each one requires web research, source validation, rule-based qualification, and structured output.
Task 1: Competitive gap analysis
The model acts as a strategy analyst comparing AIMultiple with other research and benchmarking platforms.
The agent reviews both websites and identifies offerings that the other platform provides but AIMultiple lacks. Valid categories are arena, benchmark, evaluation_format, content, methodology, dataset, and other.
The task allows up to ten gaps. Each CSV row contains the gap, its category, an explanation, a live page proving that the competing offering exists, the AIMultiple pages checked, and a priority score.
When AIMultiple provides a related offering, the model must name it and explain the remaining difference. A broad category overlap alone does not establish equivalence.
How we scored the task
The hidden answer key contains ten verified gaps and a set of known false absence claims. Each supported gap adds one to A_found, while each incorrect gap adds one to B_wrong. The score is: MAX(0, A_found − B_wrong) / 10
A fabricated, unreachable, or unsupported URL removes credit for the related gap and adds a precision penalty. This design makes speculative list expansion costly, since adding a plausible claim can lower the final score.
The output must also pass run-level structural gates. The CSV needs the correct filename and columns, between 1 and 10 rows, accepted category values, valid domains, atomic cells, and a descending order of priority. Failure on any gate makes the submission invalid.
Task 2: Best-fit account-list preparation
The model acts as a growth marketer preparing an account-based marketing list for AIMultiple.
The work begins with a coverage map of AIMultiple’s live pages. The agent records category pages, benchmarks, and vendor-comparison pages, as well as vendors mentioned in article bodies and in sponsor disclosures.
It then prepares a prioritized list of 100 accounts. Each company must be a B2B technology vendor with annual revenue between $100 million and $1 billion, headquartered in the US, Europe, or Israel, and active in a category covered by AIMultiple. The company also needs a current growth-marketing signal and at least one AIMultiple page where it represents a genuine commercial opportunity.
The account file records the company’s domain, LinkedIn page, headquarters, category, revenue and source, employee count, founding year, funding or ownership status, relevant AIMultiple pages, marketing signal, and fit score.
Parent companies and subsidiaries cannot both appear for the same vendor opportunity. The rules also exclude service firms, consumer businesses, analyst or review companies, and candidates who do not meet the revenue or headquarters requirements.
How we scored the task
The rubric contains 71 points across seven sections. The coverage-map section checks whether the model reviewed at least 30 live AIMultiple pages and included every required page type. Reviewers sample rows to confirm that named vendors and sponsor disclosures match the live pages.
The account-file section requires exactly 100 distinct domains, the specified columns, accepted classifications, descending fit scores, and complete evidence fields. A fixed sample of 20 accounts is then used for detailed checks covering revenue, LinkedIn identity, true headquarters, category fit, growth-marketing evidence, and B2B vendor status.
Other sections test whether proposed pages are genuine openings, penalize disqualified or fabricated accounts, reward the inclusion of verified high-fit companies, and check whether models avoid known near-misses.
A missing account file, incorrect columns, or any row count other than 100 makes the run invalid.
Agentic marketing benchmark methodology
We built both benchmarks from workflows used by our strategy and growth teams.
The competitive-analysis task supports decisions about product coverage, content priorities, and positioning. The account-list task supports account-based marketing and sales research.
For each workflow, we created a task file that defines:
- the model’s role
- available tools
- evidence requirements
- qualification rules
- required outputs
- time limit
We designed the tasks around our own pages and commercial workflows, and our data forms part of the answer key. This makes the evidence easier to verify, though it also limits how broadly the results apply to other companies and marketing environments.
Task instructions and scoring rubrics
For competitive analysis, models could research the live web through the Bright Data MCP. We disabled code execution and set a 90-minute limit.
For account-list preparation, models used the same web access, with code execution enabled. We allowed up to 240 minutes because the task required two output files and 100 qualified accounts.
We gave each model the task instructions and output schema, while keeping the scoring rubric hidden.
Models had to use current sources
Both tasks depend on information that changes in time, including website offerings, benchmark inventories, product pages, company revenue, headquarters, ownership, sponsor disclosures, and marketing activity.
We counted a URL when it resolved and supported the attached claim. A company homepage could not support a specific revenue figure, and a generic marketing database could not prove that a platform offered a particular benchmark format. We last verified the answer keys on July 7, 2026.
Outputs were designed for operational use
Both tasks require CSV files with fixed columns and atomic cells. This format allows us to automatically check for missing fields, duplicates, invalid categories, unsupported domains, incorrect row counts, malformed values, and sorting errors.
It also reflects how marketing teams use these outputs. Account lists and competitive analyses often move into spreadsheets, CRM workflows, internal databases, or automation tools. A structurally invalid file can block the next step even when parts of the research are correct.
We report single-run scores without variance
The results include one reported score per model. We do not currently report repeated runs, seeds, confidence intervals, or error bars.
Readers should therefore treat small score differences cautiously. In competitive analysis, one additional net correct gap changes the score by ten points. In account-list preparation, a few sampled rows can determine whether an entire threshold-based criterion passes.
We combined deterministic checks with an LLM judge
Deterministic scripts handled everything software can settle directly: file existence, CSV parsing, column order, row counts, accepted values, URL domains, duplicates, and sorting. Scripts also fetched every cited URL, followed redirects, and passed the resolved page content to the judge.
An LLM judge scored the semantic criteria, whether two offerings are equivalent, whether an absence claim holds, whether a company fits a category, whether a fetched page supports the claim attached to it, and whether a proposed page is a genuine opening.
The judge scored one row or one gap at a time against the rubric’s pass-or-fail rules, and scripts aggregated those judgments into the threshold criteria (for example, ≥18 of 20 sampled rows). It never assigned a general quality rating. Our own analysts built and froze the answer key: the gold accounts, the traps, and the sponsor-disclosure transcriptions, on July 7, 2026; the judge applied that key rather than forming its own view of what should count.
This allowed us to score two long, evidence-heavy files, but it carries a risk: a judge can accept a source it read too generously, which is the same failure mode the rubric penalizes in the models. These results use a single judging pass, with no multi-judge agreement check and no calibration against a human-scored sample.
FAQs
Agentic marketing involves using AI-powered agents to plan and complete multi-step marketing tasks with minimal human input. These agents use customer data, business rules, real-time context, and natural-language instructions to make decisions throughout marketing workflows.
Traditional marketing automation executes predefined rules, such as sending an email after a form submission or moving a lead between audience segments. Agentic systems can interpret a business goal, select the required tools, complete campaign setup, monitor campaign performance, and adjust actions to optimize performance.
An agentic marketing platform can support tasks such as:
– analyzing customer behavior and identifying audience segments;
– adapting customer journeys based on new signals;
– allocating ad spend across marketing campaigns;
– generating content within brand guidelines;
– coordinating creative workflows and preserving creative direction;
– producing AI-generated insights for journey optimization;
– moving data between disconnected tools in a marketing cloud;
– assigning specific tasks to multiple agents.
Competitive gap analysis requires evidence from both sides of every finding. The model must confirm that the competing platform provides a specific offering. It must then inspect AIMultiple’s relevant pages and determine whether an equivalent exists.
The second step requires wider site exploration. Relevant evidence may appear on a benchmark page, a category page, an article, a tool page, or a sitemap entry.
Semantic similarity adds another challenge. Two offerings can address the same broad market while using different evaluation formats or methodologies. The agent must decide whether the difference creates a meaningful gap.
The rubric also penalizes speculative findings. Broad search behavior can improve recall and increase the number of false absence claims. Conservative behavior can protect precision and miss expected gaps.
The two tasks test different combinations of capabilities. Competitive analysis requires comparative research, semantic judgment, and careful verification of absence.
Account-list preparation requires sustained web research, repeated qualification, entity resolution, source management, and exact output control.
A model can perform well on a short comparison task but fail to complete a large, structured file. Another model can handle repeated account research and fail a structural gate in competitive analysis.
Model selection for agentic marketing should therefore consider the specific marketing workflows a team plans to automate.
Cite this research
Pick the format that matches where you're publishing. Pasting the link version into your CMS preserves the backlink.
@misc{ermut2026,
author = {Ermut, Sıla and Kalelioğlu, Berk},
title = {{AIM Agentic Marketing Benchmark}},
year = {2026},
month = jul,
howpublished = {\url{https://aimultiple.com/agentic-marketing}},
note = {AIMultiple. Retrieved July 17, 2026}
}
Be the first to comment
Your email address will not be published. All fields are required. Comments are left in their original language.