Premium
Services
Premium

Web Data Scraping Benchmarks

Provider success rates, response times and metadata coverage across our web scraping benchmarks. See the methodology.

Top provider
Bright Data
Highest success rate across all benchmarks
Lowest effective cost
Nimble
$1.00 per 1k successful pages
Fastest
SerpApi
~1.1 s average completion time
Most metadata fields
Bright Data
About 28 average fields across all benchmarks
Average Success Rate of Web Data Providers
*Average success rate across all the benchmarks we ran.

Leaderboard

Filter by tool type and must-have requirements. Sort any column.

#
Model
Index
E-commerce
Social media
Search engines
Video & streaming
Travel
Reviews
Food delivery
App stores
Real estate
Job postings
News & publishing
Developer & SaaS
Finance
Government & education
1
Bright Data
Bright Data
All tool types
86.7
81.470.460.698.698.97894.399.888.590.697.396.391.591.8
2
Oxylabs
Oxylabs
All tool types
77.2
19.870.4099.789.156.290.899.671.177.5----
3
Apify
Apify
All tool types
75.4
18--99.699.2-55.7-------
4
Nimble
Nimble
All tool types
74
71.764097.719.352.484.999.78669.699.699.5-97
5
Zyte
Zyte
All tool types
72.5
25.5-699.696.36582.299.565.657.897.698.6-96.7
6
Decodo
Decodo
All tool types
69.5
39.670.1099.583.235.99099.774.877.4----

Rank#1
E-commerce81.4
Social media70.4
E-commerce
81.4
Social media
70.4
Search engines
60.6
Video & streaming
98.6
Travel
98.9
Reviews
78
Food delivery
94.3
App stores
99.8
Real estate
88.5
Job postings
90.6
News & publishing
97.3
Developer & SaaS
96.3
Finance
91.5
Government & education
91.8

Rank#2
E-commerce19.8
Social media70.4
E-commerce
19.8
Social media
70.4
Search engines
0
Video & streaming
99.7
Travel
89.1
Reviews
56.2
Food delivery
90.8
App stores
99.6
Real estate
71.1
Job postings
77.5

Rank#3
E-commerce18
E-commerce
18
Video & streaming
99.6
Travel
99.2
Food delivery
55.7

Rank#4
E-commerce71.7
Social media64
E-commerce
71.7
Social media
64
Search engines
0
Video & streaming
97.7
Travel
19.3
Reviews
52.4
Food delivery
84.9
App stores
99.7
Real estate
86
Job postings
69.6
News & publishing
99.6
Developer & SaaS
99.5
Government & education
97

Rank#5
E-commerce25.5
E-commerce
25.5
Search engines
6
Video & streaming
99.6
Travel
96.3
Reviews
65
Food delivery
82.2
App stores
99.5
Real estate
65.6
Job postings
57.8
News & publishing
97.6
Developer & SaaS
98.6
Government & education
96.7

Rank#6
E-commerce39.6
Social media70.1
E-commerce
39.6
Social media
70.1
Search engines
0
Video & streaming
99.5
Travel
83.2
Reviews
35.9
Food delivery
90
App stores
99.7
Real estate
74.8
Job postings
77.4

Cost vs. performance

Web scraper cost across our 100 e-commerce domain benchmark, and web unblocker cost across 10,000 domains. Volume-adjusted per successful request.

Pricing scenario

Avg is the effective per-1k rate we paid in our benchmark, volume-adjusted with each provider's cheapest available plan. Min is the cheapest tier each provider offers. Max is the most expensive tier.

Provider profile

Price per successful page, completion time, metadata depth, domains returning structured JSON and AI readiness for each provider.

#
Model
Domains with JSON
Avg metadata fields
$ per 1k
Completion time (s)
AI ready
1
Bright Data
Bright Data
All tool types
24281.51879
2
Oxylabs
Oxylabs
All tool types
124214.271
3
Apify
Apify
All tool types
104213.726.755
4
Nimble
Nimble
All tool types
60112.668
5
Zyte
Zyte
All tool types
00313.660
6
Decodo
Decodo
All tool types
602.416.564
7
SerpApi
SerpApi
All tool types
1328-1.1-

Cost1.5
Latency18
Context-
Domains with JSON
24
Avg metadata fields
28
$ per 1k
1.5
Completion time (s)
18
AI ready
79

Cost2
Latency14.2
Context-
Domains with JSON
12
Avg metadata fields
4
$ per 1k
2
Completion time (s)
14.2
AI ready
71

Cost13.7
Latency26.7
Context-
Domains with JSON
10
Avg metadata fields
42
$ per 1k
13.7
Completion time (s)
26.7
AI ready
55

Cost1
Latency12.6
Context-
Domains with JSON
6
Avg metadata fields
0
$ per 1k
1
Completion time (s)
12.6
AI ready
68

Cost3
Latency13.6
Context-
Domains with JSON
0
Avg metadata fields
0
$ per 1k
3
Completion time (s)
13.6
AI ready
60

Cost2.4
Latency16.5
Context-
Domains with JSON
6
Avg metadata fields
0
$ per 1k
2.4
Completion time (s)
16.5
AI ready
64

Cost-
Latency1.1
Context-
Domains with JSON
13
Avg metadata fields
28
Completion time (s)
1.1

Behaviour under load

Success rate tracked at 1, 100, 500 and 5,000 parallel requests.

Filter by concurrency level
*Providers compared by success rate based on correct text presence on extracted page and completion time for successful requests (in seconds), across different concurrency levels.

Anti-bot landscape

Which vendor guards your targets predicts your success rate.

Web unblocks across the top 10k domains

Anti-bot vendor
*Whiskers are the 95% confidence interval.

Scraping APIs across the top 10k domains

Anti-bot vendor
*Whiskers are the 95% confidence interval.

Domain coverage by provider

Which of the Tranco top 100 domains each provider can scrape, and for which of them it also returns structured JSON. Scraping support comes from our web unblocker benchmark; JSON support comes from a separate test of parsed output on the same domains. Domains are sorted by coverage, so the ones most providers support appear at the top.

Domain
google.com
chatgpt.com
bing.com
amazon.com
facebook.com
tiktok.com
youtube.com
baidu.com
googlevideo.com
yandex.ru
apple.com
github.com
instagram.com
microsoft.com
office.com
twitter.com
wikipedia.org
yahoo.com
youtu.be
zoom.us
Domains scraped6812673671013

✓ The provider scrapes the domain. ✅ The provider scrapes the domain and returns structured JSON. ✕ The provider does not scrape the domain. Of the providers in this table, only Bright Data, Nimble and Zyte ran in the web unblocker benchmark; for the other providers the table shows JSON support only.

Web data providers markdown output performance

Markdown extraction success rate and completion time from our web unblocker benchmark, and the output formats each provider returns.

*Providers compared by success rate based on correct text presence on extracted page and completion time for successful requests (in seconds), for the markdown extraction test.
ProviderMarkdownHTML
Bright Data
Bright Data
Decodo
Decodo
Zyte
Zyte
Oxylabs
Oxylabs
Apify
Apify
Nimble
Nimble
Firecrawl
Firecrawl
Exa
Exa

Success rate by page type and concurrency

Product pages and search pages behave differently on the same domain.

Concurrency
*Success rate on product (detail) vs search (listing) pages, per provider.

Frequently asked questions

DataDome, Kasada and aggressive Akamai defenses need a provider that simply gets through. Success rate dominates everything else here - look at the Index column before any other number on the table.

Millions of predictable product pages make effective cost per successful page the binding constraint, not raw success rate. Check the $ per 1k column.

Feeding clean markdown into a context window depends on token cleanliness and latency more than metadata depth. Look at the AI ready column, then Completion time.

Instagram, LinkedIn, TikTok and X make coverage the binding constraint, not speed. Check the Domains with JSON column before anything else.

Thousands of parallel requests in short windows need a provider that holds up under load. Look at the Behaviour under load chart further down the page, not the Index column.

Procurement, KYC and audit requirements make compliance the first filter, before performance is even considered. Check a provider's Free / PAYG / paid-on-success flags and the ethical web data benchmark before the leaderboard.

Methodology

The index re-uses the raw results of every web data benchmark we published. A benchmark counts if it tests web data providers and records the outcome of each request. A provider's index score is the average of its success rates across the benchmarks it ran in. Every benchmark carries equal weight, so a 1,400-request study counts as much as a 260,000-request one. Success means the response returned the target page with the content we asked for. We check that against bot pages and empty shells, so a 200 response with a challenge page counts as a failure. If a provider was never tested in a benchmark, we show a dash instead of a zero.

Targets come from the Tranco list, which ranks domains by averaging several traffic rankings. We remove dead, adult, gambling and malicious hosts before testing. Domains are grouped into categories: e-commerce, social media, search engines, video and streaming, travel, reviews, food delivery, app stores, real estate, job postings, news and publishing, developer and SaaS, finance, and government and education. Domain coverage and structured JSON support are checked on the top 100. Cost is per 1,000 successful pages, not per 1,000 requests sent. We take the cheapest plan a provider offers at that volume and divide by the success rate we measured. Completion time is how long one request takes from start to finish. Metadata fields are the parsed fields a provider returns when it answers with JSON. AI ready is the share of responses that came back as clean markdown. The anti-bot results come from two of our runs, the unblocker benchmark across 10,000 domains and the e-commerce benchmark across the top 100 domains, with each domain labelled by the vendor guarding it. Whiskers are 95% confidence intervals.

Explore Web Data Scraping Benchmarks

Is Web Scraping Legal? Laws & Best Practices

Web Data Scraping
Insight
Sep 17

Legal regulations have changed in the web scraping market. While litigation once focused on unauthorized access, new lawsuits related to AI training and technical workarounds are shaping acceptable practices. Disclaimer: Our work is for informational purposes and not legal advice; please get professional legal advice for specific guidance. Web scraping is legal if you scrape…

Read More
Web Data Scraping
Benchmark
Sep 15

Ethical & Compliant Web Data Benchmark

As enterprises scale their web data operations, compliance, data, and risk executives increasingly evaluate the associated ethical, reputational, and legal risks. We benchmarked 5 leading web data collection services across 3 dimensions and tested each service with more than 20 potentially unethical scenarios. Our work helps you assess the ethical standing of your data collection…

Web Data Scraping
Benchmark
Sep 10

Top 5 Job Posting Scraper APIs Compared

We benchmarked 5 leading web scraping providers across 5 major job platforms by running 12,500 requests in total, then measured each provider’s success rate, completion time, and metadata output. You can read benchmark methodology section for more details on the testing process = supported, returns HTML = supported, returns structured data = no data returned…

Web Data Scraping
Open World Evaluation
Sep 10

Top 5 Home Depot Scrapers Benchmarked & Compared

We benchmarked five web data providers on Home Depot, each fetching the same 50 product and search pages at 5 concurrent requests, for a total of 250 requests. You can read more about our benchmark methodology. Bright Data offers a dedicated scraper API for Home Depot, Apify provides a general e-commerce actor, and SerpApi also…

Web Data Scraping
Feature Comparison
Sep 10

Best ScrapeBox Alternatives

ScrapeBox is a Windows and macOS desktop application used for SEO tasks such as search engine scraping, keyword harvesting, link building, comment posting, and backlink checking. However, it is a desktop GUI tool, not an API, and the cost is higher once premium plugins, proxies, and a CAPTCHA service are added. So alternatives are grouped…

Web Data Scraping
Open World Evaluation
Sep 10

eBay Scraping: Top 6 Providers Compared

We benchmarked eBay across 4 web scraping providers, totaling 1,400 requests over search and product pages of 7 eBay country sites, measuring both success rate and end-to-end completion time. You can read more about our benchmark methodology. Bright Data, Apify and SerpApi were the three providers that returned parsed JSON. Field counts came to 59…

Web Data Scraping
Open World Evaluation
Sep 7

Top Firecrawl Alternatives: Features & Pricing

Firecrawl is an AI-native web scraping and crawling API that converts dynamic web pages into LLM-ready Markdown and structured data. We compared its alternatives across benchmarks and features. We left out JavaScript rendering, search endpoints, structured JSON output and MCP server support. Nearly every product here offers all four, so they no longer separate one…

Web Data Scraping
Benchmark
Aug 21

Top 7 Video Scrapers: Tested & Ranked

We tested the top 7 video scraping providers to see how they handle video metadata on the top video platform, totaling 6,000 requests, and measured their success rate, response time, and metadata fields. To see how we calculated these metrics, read video scraping benchmark methodology. Different providers return different amounts of metadata for the same…

Web Data Scraping
Benchmark
Aug 21

Top 6 LLM Scrapers: ChatGPT, Perplexity & Gemini

We benchmarked how the top LLM scraper providers, including Bright Data, Oxylabs, and Apify, perform at extracting outputs from LLM platforms such as ChatGPT, Gemini, Perplexity, and Google AI Mode. To ensure reliable results, we ran 1,000 tests per provider, repeating each prompt 10 times for consistency. The top-performing provider is detailed below. Providers missing…

Web Data Scraping
Insight
Aug 14

How to Bypass CAPTCHA (reCAPTCHA & hCaptcha)

Modern CAPTCHA and human-verification systems use a mix of challenge-response tests, browser signals, server-side token validation, and adaptive challenges. Attempting to bypass CAPTCHA on third-party websites can violate the terms of service or trigger account or IP blocks. The better approach is to use official APIs, reduce request rates, or implement a modern bot-management solution…

Web Data Scraping
Insight
Aug 13

The Most Common Web Scraping Challenges

Web scraping has become more difficult in recent years. Since 2025, AI-related scraping has raised significant legal concerns. Platforms and infrastructure providers have adopted new methods to control AI crawlers and manage data collection. There are many technical challenges that web scrapers face due to the barriers set by data owners or website owners to…