We are introducing the AIM Agentic Marketing Benchmark, which measures agent performance on three marketing workflows: competitive gap analysis, ABM target list preparation, and a personalized sales deck.
We also ran a separate website reputation audit in which agents examined AIMultiple’s English-language content and reported verifiable issues with factual accuracy, citations, consistency, freshness, functionality, grammar, and formatting that could undermine readers’ trust.
Agentic marketing benchmark results
The task scores are normalized to a 0–100 scale.
- For competitive analysis, the score equals: 100 × MAX(0, verified gaps found − incorrect gaps) / 10
- For account-list preparation, the score is the percentage of 71 available rubric points earned by the model.
- For the sales deck task, the score is the sum of 5 20-point sections.
- The overall score is the arithmetic mean of the competitive-analysis, account-list, and sales-deck scores. The website reputation audit does not contribute to this overall score. Its raw score is the sum of the priority values of all verified issues.
Unlike the three normalized tasks, the audit has no fixed maximum score. Therefore, we included the results in a separate evaluation and did not combine them with the overall benchmark.
Read the methodology for more information on task evaluation.
The real-life tasks used in the benchmark
These agentic marketing workflows cover strategy and content coverage, growth, sales research, personalized content creation, and website quality assurance. Each requires live research, source validation, rule-based judgment, and an output that can be integrated into an operational workflow.
Task 1: Competitive gap analysis
The model works as a strategy analyst at AIMultiple and compares it against one competing AI benchmarking platform. The run allows 90 minutes of runtime and live web access via the Bright Data MCP, with code execution disabled.
The output is a single file containing up to 10 rows, ordered by priority. Each row names an offering the competitor publishes but AIMultiple does not, explains the difference in one or two sentences, cites the competitor page the model opened, and lists the AIMultiple pages it examined before claiming the absence.
The absence claim is the harder half of each finding. Identifying a feature on the competitor’s site requires a single page. Establishing that AIMultiple does not provide it requires reviewing the category pages, benchmarks, articles, and sitemap.
Where AIMultiple publishes an adjacent offering, the model must name that page and specify what the competitor’s version does differently. Coverage of the same broad subject does not establish an equivalent offering.
How we scored the task
Scoring runs against an answer key compiled by hand: ten gaps verified on both sites, together with the absence claims that live AIMultiple pages disprove.
Each verified gap adds a point, and each incorrect absence claim deducts one. An incorrect claim includes any row whose URL is fabricated, unreachable, or does not support the attached claim.
The key holds ten gaps, but a model can find a real one that is not on it. A reviewer reads those findings by hand and gives the point if the cited page shows the offering, no live AIMultiple page contradicts the claim, and the offering is truly missing from AIMultiple.
Two rules limit these extra points:
- A finding that repeats one of the existing ten in different words earns nothing extra. It is credited under the listed gap instead.
- A finding that fails review costs a point. If AIMultiple publishes the offering, or the competitor does not have it, the row scores zero, the same penalty as any other wrong claim.
This prevents the key from penalizing a finding the analysts did not record, while withholding credit from speculative claims.
As the last part of the evaluation, the reviewer performs link checks. Link checks follow redirects and evaluate the final destination, so a 301 does not fail a row on its own. An incorrect filename, missing or reordered columns, a row count outside one to ten, a priority outside 1 to 10, an evidence URL on the wrong domain, or rows out of order each render the run invalid, and an invalid run receives no numerical score.
Task 2: Best-fit account-list preparation
The model acts as a growth marketer preparing an account-based marketing list for AIMultiple.
The work begins with a coverage map of AIMultiple’s live pages. The agent records category pages, benchmarks, and vendor-comparison pages, as well as vendors mentioned in article bodies and in sponsor disclosures.
It then prepares a prioritized list of 100 accounts. Each company must be a B2B technology vendor with annual revenue between $100 million and $1 billion, headquartered in the US, Europe, or Israel, and active in a category covered by AIMultiple. The company also needs a current growth-marketing signal and at least one AIMultiple page where it represents a genuine commercial opportunity.
The account file records the company’s domain, LinkedIn page, headquarters, category, revenue and source, employee count, founding year, funding or ownership status, relevant AIMultiple pages, marketing signal, and fit score.
Parent companies and subsidiaries cannot both appear for the same vendor opportunity. The rules also exclude service firms, consumer businesses, analyst or review companies, and candidates who do not meet the revenue or headquarters requirements.
How we scored the task
The rubric contains 71 points across seven sections. The coverage-map section checks whether the model reviewed at least 30 live AIMultiple pages and included every required page type. Reviewers sample rows to confirm that named vendors and sponsor disclosures match the live pages.
The account-file section requires exactly 100 distinct domains, the specified columns, accepted classifications, descending fit scores, and complete evidence fields. A fixed sample of 20 accounts is then used for detailed checks covering revenue, LinkedIn identity, true headquarters, category fit, growth-marketing evidence, and B2B vendor status.
Other sections test whether proposed pages are genuine openings, penalize disqualified or fabricated accounts, reward the inclusion of verified high-fit companies, and check whether models avoid known near-misses.
A missing account file, incorrect columns, or any row count other than 100 makes the run invalid.
Task 3: Personalized sales deck
Note: We excluded Kimi K3 and Gemini 3.5 Flash because each requested access to a directory outside its working folder; the harness declined the request automatically, and both runs ended there. Neither produced a deck, and neither failed the task, as nothing they produced was judged. A zero would put an infrastructure problem on the model’s record, so both are marked as not run and excluded from the average.
The model serves as a pre-sales analyst at AIMultiple and builds a pitch deck for a senior product-marketing executive. The run includes 90 minutes, live web access via the Bright Data MCP, code execution, and a folder of brand assets containing the logos and fonts.
Most of the research goes into the executive rather than the company. This executive runs marketing for a single product line and has publicly stated what counts when judging hardware: how many tokens a dollar buys, whether an independent third party checked the numbers, and whether the test resembled a real workload or a demo.
Another recurring argument is that a rack is a fairer unit of comparison than a chip. As these arguments are public and can be found in interviews, conference talks, and posts, the task is to read them and then find the AIMultiple benchmarks that answer them. A deck can be accurate about AIMultiple and still miss the person entirely, and that is the version that scores poorly.
Here are some outputs from Fable 5, Gemini 3-1 pro-preview, and GPT 5.6 Sol:
How we scored the task
The deck is scored on five aspects, each worth 20 points: whether it follows AIMultiple’s visual identity, whether it cites the right benchmarks, whether it proposes useful next steps, whether it explains what AIMultiple is worth to the recipient, and whether the slides are clear. Each of the five breaks into small checks with fixed answers.
Visual identity is the one part a model can get right without doing any research. The logo files, the hex value of the accent color, the text colors, the typeface, and the spelling of the company name are all either correct or not. For example, a deck set in Calibri loses the points.
The second part evaluates what the deck points to today. AIMultiple publishes many articles on hardware, and some of them test the kind of chips this person sells. Therefore, the points go to a deck citing data-center GPU studies rather than whichever hardware pages turned up first.
The third section is about the work AIMultiple could do next. A suggestion counts here if it names the study and the number it would produce. For example, re-running an existing comparison on this year’s chips or adding the company’s hardware to a benchmark that currently excludes it earns a point. On the other hand, proposing that the two firms look at inference performance together names neither a study nor a measurement, therefore earns no points.
The fourth part concerns what AIMultiple is worth to the recipient. Points are awarded for three statements that:
- AIMultiple is independent
- Its readership is of a stated size supported by a named source
- This readership serves a specific purpose for the recipient’s business, such as reaching buyers at the shortlist stage
A deck that lists what AIMultiple publishes without connecting any of it to the recipient’s own products does not earn points.
The fifth part covers whether the deck can be read. Two checks matter more than the rest: headlines have to state a takeaway rather than label the slide, and numbers have to say where they came from, with the company’s figures kept apart from AIMultiple’s own measurements. The rest is mechanical, such as evaluating the slide count, text overflow, chart labels, filler.
Four of the five parts also carry penalties. An invented benchmark, a made-up figure, a dead link, or presenting the company’s own marketing claim as an AIMultiple measurement costs points. Promising favorable results, preferential ranking, or a review before publication also cost points, since AIMultiple offers none of them. The deck has to persuade without fabricating evidence and without offering anything the business does not sell.
Task 4: Website reputation audit
The model acts as a website quality and reputation auditor for AIMultiple. The run allows 120 minutes, live web access through the Bright Data MCP, and code execution.
The agent begins by fetching AIMultiple’s post sitemap and excluding URLs with German, Spanish, French, Italian, Portuguese, or Turkish language prefixes. It must examine at least 100 remaining English-language URLs.
For each page, the agent searches for issues that could damage AIMultiple’s credibility. These can include inconsistent information across pages, broken links, outdated claims, factual errors, poor grammar or formatting, missing citations, misleading claims, broken functionality, and generic AI-generated filler.
The output is a Results.csv file containing no more than 50 rows. Each row must describe one distinct, deduplicated issue and aggregate all URLs where that issue occurs.
Each row records:
- the issue ID;
- priority;
- severity;
- number of impacted URLs;
- impacted URLs;
- exact quoted evidence;
- an explanation of the reputational problem;
- and a recommended fix.
The file must be ordered by priority from highest to lowest.
How we scored the task
Priority is calculated as: Priority = Severity × Number of impacted URLs
Severity uses four levels:
- 4: Factual errors, broken functionality or other issues that can cause significant reputational harm.
- 3: Cross-page inconsistencies, outdated information, missing sources, broken links or AI-style filler.
- 2: Major grammar or formatting problems.
- 1: Minor grammar, formatting or style problems that do not affect readability.
A row contributes to the model’s score only when all applicable checks pass. Deterministic scripts first check the CSV structure, column names, encoding, priority arithmetic, severity value, URL counts, URL language and domain, and whether quoted evidence appears on the attributed page.
A live web judge then checks whether the issue reproduces on the sampled URLs, whether the evidence demonstrates the described defect, whether the assigned severity is appropriate, and whether the finding represents a real problem (not a false positive or an editorial preference).
Lastly, the complete submission is checked for duplicate findings and issues that are too broad to verify. For example, “citation issues” is too general for one row, while “broken citation links” is a sufficiently specific issue.
The final raw score is the sum of priority across all rows that pass every check. A model that submits an invalid file or no valid data rows receives a score of zero. See how each model performed below:
Agentic marketing benchmark methodology
We built the original benchmark from three workflows used by our strategy, growth, and sales teams. We then added a website reputation audit representing content quality assurance and editorial governance.
The competitive-analysis task supports decisions about product coverage, content priorities, and positioning. The account-list task supports account-based marketing and sales research. The sales deck supports pre-sales outreach: selecting a prospect, determining which published benchmarks matter, and building the document.
The website reputation audit supports content maintenance and trust. It identifies broken evidence, stale information, cross-page inconsistencies, factual problems, functionality failures, and editorial defects.
For each workflow, we created a task file that defines:
- the model’s role
- available tools
- evidence requirements
- qualification rules
- required outputs
- time limit
We designed the tasks around our own pages and commercial workflows, and our data forms part of the answer key. This makes the evidence easier to verify, though it also limits how broadly the results apply to other companies and marketing environments.
Task instructions and scoring rubrics
For competitive analysis, models could research the live web through the Bright Data MCP. We disabled code execution and set a 90-minute limit.
For account-list preparation, models used the same web access, with code execution enabled. We allowed up to 240 minutes because the task required two output files and 100 qualified accounts.
For sales deck preparation, we enabled the same web access with code execution. The duration was 90 minutes, with a folder of brand assets in the working directory.
For the website reputation audit, models used live web access through the Bright Data MCP with code execution enabled. We allowed up to 120 minutes as the task required crawling and analyzing at least 100 pages. The required output was a single CSV containing no more than 50 deduplicated issue groups.
We gave each model the task instructions and output schema, while keeping the scoring rubric hidden.
Models had to use current sources
All four workflows depend on information that can change. This includes site offerings, benchmark inventories, product pages, company revenue, headquarters, ownership, sponsor disclosures, marketing activity, public statements, article content, citations, and link destinations.
We counted a URL only when it resolved and supported the claim attached to it. A company homepage cannot support a specific revenue figure, and a page containing a similar phrase cannot support an audit finding unless the submitted quotation and described defect are present.
The answer keys for the competitive-analysis and account-list tasks were last verified on July 7, 2026. The sales-deck rubric was verified on July 25, 2026, and the decks were scored the following day.
The website reputation audit did not use a fixed list of expected issues. Each submitted finding was checked against AIMultiple’s live pages at the time of scoring. This is necessary because the website is continuously updated, and an issue that existed during one run may be fixed later.
Outputs were designed for operational use
The competitive-analysis, account-list and website-reputation tasks require CSV files with fixed columns and atomic cells. This allows us to automatically check for missing fields, duplicates, unsupported domains, incorrect row counts, malformed values, arithmetic errors and sorting problems.
The sales-deck task requires a 16:9 PowerPoint file that opens without a repair prompt.
Competitive analyses and account lists can move into spreadsheets, CRM workflows, and automation tools. The reputation audit output can be moved into an editorial or engineering issue tracker. The deck can be sent to a prospect after review.
A structurally invalid file blocks the next operational step even when some of the research inside it is useful. For this reason, file delivery and schema compliance are part of the benchmark rather than administrative requirements outside it.
We report single-run scores without variance
The results include one reported score per model. We do not currently report repeated runs, seeds, confidence intervals, or error bars.
Readers should therefore treat small score differences cautiously. In competitive analysis, one additional net correct gap changes the score by ten points. In account-list preparation, a few sampled rows can determine whether an entire threshold-based criterion passes.
We combined deterministic checks with an LLM judge
Deterministic scripts handled everything software can settle directly: file existence, CSV parsing, column order, row counts, accepted values, URL domains, duplicates, sorting, and for the deck, archive integrity, slide dimensions, and font declarations. Scripts also fetched every cited URL, followed redirects, and passed the resolved page content to the judge.
An LLM judge scored the semantic criteria one row, gap, or criterion at a time against pass-or-fail rules, and never assigned a general quality rating. Our analysts built the answer keys, and the judge applied them.
Competitive analysis adds one human step. A model can report a real gap that the key does not list, so those findings go to an analyst, who credits them after checking both sites.
Cite this research
Pick the format that matches where you're publishing. Pasting the link version into your CMS preserves the backlink.
@misc{ermut2026,
author = {Ermut, Sıla and Kalelioğlu, Berk},
title = {{AIM Agentic Marketing Benchmark}},
year = {2026},
month = aug,
howpublished = {\url{https://aimultiple.com/agentic-marketing}},
note = {AIMultiple. Retrieved August 5, 2026}
}
Be the first to comment
Your email address will not be published. All fields are required. Comments are left in their original language.