Premium
Services
Premium

Web Data Scraping Benchmarks

Provider success rates, response times and metadata coverage across our web scraping benchmarks.

Top provider
Bright Data
Highest success rate across all benchmarks
Lowest effective cost
Nimble
$1.00 per 1k successful pages
Fastest
SerpApi
~1.1 s average completion time
Most metadata fields
Bright Data
About 28 average fields across all benchmarks
Average Success Rate of Web Data Providers
*Average success rate across all the benchmarks we ran.

Leaderboard

Filter by tool type and must-have requirements. Sort any column.

#
Model
Index
E-commerce
Social media
Search engines
Video & streaming
Travel
Reviews
Food delivery
App stores
Real estate
Job postings
News & publishing
Developer & SaaS
Finance
Government & education
Other
1
Bright Data
Bright Data
All tool types
86.7
81.470.460.698.598.97864.199.888.590.697.396.391.591.896.8
2
Oxylabs
Oxylabs
All tool types
77.2
19.870.4099.789.156.291.799.671.177.5-----
3
Apify
Apify
All tool types
75.4
18--99.699.298.644.6-94------
4
Nimble
Nimble
All tool types
74
71.764097.719.352.46899.78669.699.699.5-9796.9
5
Zyte
Zyte
All tool types
72.5
25.5-699.696.36575.999.565.657.897.698.6-96.797.8
6
Decodo
Decodo
All tool types
69.5
39.670.1099.583.235.991.399.774.877.4-----

Rank#1
E-commerce81.4
Social media70.4
E-commerce
81.4
Social media
70.4
Search engines
60.6
Video & streaming
98.5
Travel
98.9
Reviews
78
Food delivery
64.1
App stores
99.8
Real estate
88.5
Job postings
90.6
News & publishing
97.3
Developer & SaaS
96.3
Finance
91.5
Government & education
91.8
Other
96.8

Rank#2
E-commerce19.8
Social media70.4
E-commerce
19.8
Social media
70.4
Search engines
0
Video & streaming
99.7
Travel
89.1
Reviews
56.2
Food delivery
91.7
App stores
99.6
Real estate
71.1
Job postings
77.5

Rank#3
E-commerce18
E-commerce
18
Video & streaming
99.6
Travel
99.2
Reviews
98.6
Food delivery
44.6
Real estate
94

Rank#4
E-commerce71.7
Social media64
E-commerce
71.7
Social media
64
Search engines
0
Video & streaming
97.7
Travel
19.3
Reviews
52.4
Food delivery
68
App stores
99.7
Real estate
86
Job postings
69.6
News & publishing
99.6
Developer & SaaS
99.5
Government & education
97
Other
96.9

Rank#5
E-commerce25.5
E-commerce
25.5
Search engines
6
Video & streaming
99.6
Travel
96.3
Reviews
65
Food delivery
75.9
App stores
99.5
Real estate
65.6
Job postings
57.8
News & publishing
97.6
Developer & SaaS
98.6
Government & education
96.7
Other
97.8

Rank#6
E-commerce39.6
Social media70.1
E-commerce
39.6
Social media
70.1
Search engines
0
Video & streaming
99.5
Travel
83.2
Reviews
35.9
Food delivery
91.3
App stores
99.7
Real estate
74.8
Job postings
77.4

Cost vs. performance

Web scraper cost across our 100 e-commerce domain benchmark, and web unblocker cost across 10,000 domains. Volume-adjusted per successful request.

Pricing scenario

Avg is the effective per-1k rate we paid in our benchmark, volume-adjusted with each provider's cheapest available plan. Min is the cheapest tier each provider offers. Max is the most expensive tier.

Provider profile

Price per successful page, completion time, metadata depth, domains returning structured JSON and AI readiness for each provider.

#
Model
Domains with JSON
Avg metadata fields
$ per 1k
Completion time (s)
AI ready
1
Bright Data
Bright Data
All tool types
24281.51879
2
Oxylabs
Oxylabs
All tool types
124214.271
3
Apify
Apify
All tool types
104213.726.755
4
Nimble
Nimble
All tool types
60112.668
5
Zyte
Zyte
All tool types
00313.660
6
Decodo
Decodo
All tool types
602.416.564
7
SerpApi
SerpApi
All tool types
1328-1.1-

Cost1.5
Latency18
Context-
Domains with JSON
24
Avg metadata fields
28
$ per 1k
1.5
Completion time (s)
18
AI ready
79

Cost2
Latency14.2
Context-
Domains with JSON
12
Avg metadata fields
4
$ per 1k
2
Completion time (s)
14.2
AI ready
71

Cost13.7
Latency26.7
Context-
Domains with JSON
10
Avg metadata fields
42
$ per 1k
13.7
Completion time (s)
26.7
AI ready
55

Cost1
Latency12.6
Context-
Domains with JSON
6
Avg metadata fields
0
$ per 1k
1
Completion time (s)
12.6
AI ready
68

Cost3
Latency13.6
Context-
Domains with JSON
0
Avg metadata fields
0
$ per 1k
3
Completion time (s)
13.6
AI ready
60

Cost2.4
Latency16.5
Context-
Domains with JSON
6
Avg metadata fields
0
$ per 1k
2.4
Completion time (s)
16.5
AI ready
64

Cost-
Latency1.1
Context-
Domains with JSON
13
Avg metadata fields
28
Completion time (s)
1.1

Behaviour under load

Success rate tracked at 1, 100, 500 and 5,000 parallel requests.

Filter by concurrency level
*Providers compared by success rate based on correct text presence on extracted page and completion time for successful requests (in seconds), across different concurrency levels.

Anti-bot landscape

Which vendor guards your targets predicts your success rate.

Web unblocks across the top 10k domains

Anti-bot vendor
*Whiskers are the 95% confidence interval.

Scraping APIs across the top 10k domains

Anti-bot vendor
*Whiskers are the 95% confidence interval.

JSON coverage

Each provider was tested against the Tranco top 100 domains for whether it returns structured JSON rather than raw markup. A double check mark means the provider returned parsed structured fields for that domain; a cross means it did not. Domains are sorted by coverage, so the ones supported by the most providers appear at the top.

Domain
Bright Data
Bright Data
Oxylabs
Oxylabs
Decodo
Decodo
Nimble
Nimble
Zyte
Zyte
Apify
Apify
SerpApi
SerpApi
google.com
chatgpt.com
facebook.com
youtube.com
amazon.com
bing.com
tiktok.com
googlevideo.com
microsoft.com
apple.com
instagram.com
fbcdn.net
twitter.com
office.com
github.com
wikipedia.org
youtu.be
yahoo.com
zoom.us
baidu.com
Domains with JSON24126601013

Across the Tranco top 100 domains, Bright Data returns structured JSON for 24, SerpApi for 13, Oxylabs for 12, Apify for 10, and Decodo and Nimble for 6 each - Zyte returned none in this sample.

Success rate by page type and concurrency

Product pages and search pages behave differently on the same domain.

Concurrency
*Success rate on product (detail) vs search (listing) pages, per provider.

Start here

No provider wins every column. Choose the constraint that actually binds you.

DataDome, Kasada and aggressive Akamai defenses need a provider that simply gets through. Success rate dominates everything else here - look at the Index column before any other number on the table.

Millions of predictable product pages make effective cost per successful page the binding constraint, not raw success rate. Check the $ per 1k column.

Feeding clean markdown into a context window depends on token cleanliness and latency more than metadata depth. Look at the AI ready column, then Completion time.

Instagram, LinkedIn, TikTok and X make coverage the binding constraint, not speed. Check the Domains with JSON column before anything else.

Thousands of parallel requests in short windows need a provider that holds up under load. Look at the Behaviour under load chart further down the page, not the Index column.

Procurement, KYC and audit requirements make compliance the first filter, before performance is even considered. Check a provider's Free / PAYG / paid-on-success flags and the ethical web data benchmark before the leaderboard.

Methodology

The index re-uses the raw results of every web data benchmark we published. A benchmark counts if it tests web data providers and records the outcome of each request. A provider's index score is the average of its success rates across the benchmarks it ran in. Every benchmark carries equal weight, so a 1,400-request study counts as much as a 260,000-request one. Success means the response returned the target page with the content we asked for. We check that against bot pages and empty shells, so a 200 response with a challenge page counts as a failure. If a provider was never tested in a benchmark, we show a dash instead of a zero.

Targets come from the Tranco list, which ranks domains by averaging several traffic rankings. We remove dead, adult, gambling and malicious hosts before testing. Domains are grouped into categories: e-commerce, social media, search engines, video and streaming, travel, reviews, food delivery, app stores, real estate, job postings, news and publishing, developer and SaaS, finance, government and education, and other. Structured JSON coverage is checked on the top 100. Cost is per 1,000 successful pages, not per 1,000 requests sent. We take the cheapest plan a provider offers at that volume and divide by the success rate we measured. Completion time is how long one request takes from start to finish. Metadata fields are the parsed fields a provider returns when it answers with JSON. AI ready is the share of responses that came back as clean markdown. The anti-bot results come from two of our runs, the unblocker benchmark across 10,000 domains and the e-commerce benchmark across the top 100 domains, with each domain labelled by the vendor guarding it. Whiskers are 95% confidence intervals.

Explore Web Data Scraping Benchmarks

Web Scraping Roadmap

Web Data Scraping
Benchmark
Sep 21

We scraped 10,000 live domains and 100 marketplaces using products from six web data infrastructure companies.We benchmarked these tools to see how well they handle enterprise web data use cases. We benchmarked 4 leading web data providers across the top 10,000 domains, running a total of 260,000 requests. Each provider was tested at multiple concurrency…

Read More
Web Data Scraping
Insight
Sep 17

Is Web Scraping Legal? Laws & Best Practices

Legal regulations have changed in the web scraping market. While litigation once focused on unauthorized access, new lawsuits related to AI training and technical workarounds are shaping acceptable practices. Disclaimer: Our work is for informational purposes and not legal advice; please get professional legal advice for specific guidance. Web scraping is legal if you scrape…

Web Data Scraping
Benchmark
Sep 15

Ethical & Compliant Web Data Benchmark

As enterprises scale their web data operations, compliance, data, and risk executives increasingly evaluate the associated ethical, reputational, and legal risks. We benchmarked 5 leading web data collection services across 3 dimensions and tested each service with more than 20 potentially unethical scenarios. Our work helps you assess the ethical standing of your data collection…

Web Data Scraping
Benchmark
Sep 10

Top 5 Job Posting Scraper APIs Compared

We benchmarked 5 leading web scraping providers across 5 major job platforms by running 12,500 requests in total, then measured each provider’s success rate, completion time, and metadata output. You can read benchmark methodology section for more details on the testing process = supported, returns HTML = supported, returns structured data = no data returned…

Web Data Scraping
Open World Evaluation
Sep 10

Top 5 Home Depot Scrapers Benchmarked & Compared

We benchmarked five web data providers on Home Depot, each fetching the same 50 product and search pages at 5 concurrent requests, for a total of 250 requests. You can read more about our benchmark methodology. Bright Data offers a dedicated scraper API for Home Depot, Apify provides a general e-commerce actor, and SerpApi also…

Web Data Scraping
Feature Comparison
Sep 10

Best ScrapeBox Alternatives

ScrapeBox is a Windows and macOS desktop application used for SEO tasks such as search engine scraping, keyword harvesting, link building, comment posting, and backlink checking. However, it is a desktop GUI tool, not an API, and the cost is higher once premium plugins, proxies, and a CAPTCHA service are added. So alternatives are grouped…

Web Data Scraping
Open World Evaluation
Sep 10

eBay Scraping: Top 6 Providers Compared

We benchmarked eBay across 4 web scraping providers, totaling 1,400 requests over search and product pages of 7 eBay country sites, measuring both success rate and end-to-end completion time. You can read more about our benchmark methodology. Bright Data, Apify and SerpApi were the three providers that returned parsed JSON. Field counts came to 59…

Web Data Scraping
Insight
Sep 10

Best AI Web Scraping Tools Benchmarked

Sites change their layout and the fields you need from a page shift over time. These changes break manually-coded scrapers. AI scrapers can be updated with simple prompts, and some can repair a saved scraper when the site changes. We benchmarked top AI web scraping tools across the top 10 e-commerce domains to see their…

Web Data Scraping
Open World Evaluation
Sep 7

Top Firecrawl Alternatives: Features & Pricing

Firecrawl is an AI-native web scraping and crawling API that converts dynamic web pages into LLM-ready Markdown and structured data. We compared its alternatives across benchmarks and features. We left out JavaScript rendering, search endpoints, structured JSON output and MCP server support. Nearly every product here offers all four, so they no longer separate one…

Web Data Scraping
Benchmark
Sep 2

The Best LinkedIn Scrapers

We benchmarked four LinkedIn scraping APIs with 4,000 requests, sending each provider the same 1,000 LinkedIn post URLs. For every request we measured success rate, completion time, and the number of parsed metadata fields returned. Read our methodology for details about the LinkedIn benchmark. This chart compares the daily success rates of LinkedIn scraper APIs…

Web Data Scraping
Insight
Aug 28

Best Facebook Scrapers: Apify, Bright Data & Decodo

Using Python and a managed Facebook scraping API lets you collect public posts, comments, likes, and shares. This tutorial demonstrates how to scrape Facebook posts by keyword and retrieve their URLs via Google search. Then it explains how to extract detailed post data using the API, along with tips for scaling the process with tools…