Web Data Scraping Benchmarks
Provider success rates, response times and metadata coverage across our web scraping benchmarks. See the methodology.
Leaderboard
Filter by tool type and must-have requirements. Sort any column.
# | Model | Index | E-commerce | Social media | Search engines | Video & streaming | Travel | Reviews | Food delivery | App stores | Real estate | Job postings | News & publishing | Developer & SaaS | Finance | Government & education |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Bright Data All tool types | 86.7 | 81.4 | 70.4 | 60.6 | 98.6 | 98.9 | 78 | 94.3 | 99.8 | 88.5 | 90.6 | 97.3 | 96.3 | 91.5 | 91.8 |
| 2 | Oxylabs All tool types | 77.2 | 19.8 | 70.4 | 0 | 99.7 | 89.1 | 56.2 | 90.8 | 99.6 | 71.1 | 77.5 | - | - | - | - |
| 3 | Apify All tool types | 75.4 | 18 | - | - | 99.6 | 99.2 | - | 55.7 | - | - | - | - | - | - | - |
| 4 | Nimble All tool types | 74 | 71.7 | 64 | 0 | 97.7 | 19.3 | 52.4 | 84.9 | 99.7 | 86 | 69.6 | 99.6 | 99.5 | - | 97 |
| 5 | Zyte All tool types | 72.5 | 25.5 | - | 6 | 99.6 | 96.3 | 65 | 82.2 | 99.5 | 65.6 | 57.8 | 97.6 | 98.6 | - | 96.7 |
| 6 | Decodo All tool types | 69.5 | 39.6 | 70.1 | 0 | 99.5 | 83.2 | 35.9 | 90 | 99.7 | 74.8 | 77.4 | - | - | - | - |
Cost vs. performance
Web scraper cost across our 100 e-commerce domain benchmark, and web unblocker cost across 10,000 domains. Volume-adjusted per successful request.
Avg is the effective per-1k rate we paid in our benchmark, volume-adjusted with each provider's cheapest available plan. Min is the cheapest tier each provider offers. Max is the most expensive tier.
Provider profile
Price per successful page, completion time, metadata depth, domains returning structured JSON and AI readiness for each provider.
# | Model | Domains with JSON | Avg metadata fields | $ per 1k | Completion time (s) | AI ready |
|---|---|---|---|---|---|---|
| 1 | Bright Data All tool types | 24 | 28 | 1.5 | 18 | 79 |
| 2 | Oxylabs All tool types | 12 | 4 | 2 | 14.2 | 71 |
| 3 | Apify All tool types | 10 | 42 | 13.7 | 26.7 | 55 |
| 4 | Nimble All tool types | 6 | 0 | 1 | 12.6 | 68 |
| 5 | Zyte All tool types | 0 | 0 | 3 | 13.6 | 60 |
| 6 | Decodo All tool types | 6 | 0 | 2.4 | 16.5 | 64 |
| 7 | SerpApi All tool types | 13 | 28 | - | 1.1 | - |
Behaviour under load
Success rate tracked at 1, 100, 500 and 5,000 parallel requests.
Anti-bot landscape
Which vendor guards your targets predicts your success rate.
Web unblocks across the top 10k domains
Scraping APIs across the top 10k domains
Domain coverage by provider
Which of the Tranco top 100 domains each provider can scrape, and for which of them it also returns structured JSON. Scraping support comes from our web unblocker benchmark; JSON support comes from a separate test of parsed output on the same domains. Domains are sorted by coverage, so the ones most providers support appear at the top.
| Domain | |||||||
|---|---|---|---|---|---|---|---|
| google.com | |||||||
| chatgpt.com | |||||||
| bing.com | |||||||
| amazon.com | |||||||
| facebook.com | |||||||
| tiktok.com | |||||||
| youtube.com | |||||||
| baidu.com | |||||||
| googlevideo.com | |||||||
| yandex.ru | |||||||
| apple.com | |||||||
| github.com | |||||||
| instagram.com | |||||||
| microsoft.com | |||||||
| office.com | |||||||
| twitter.com | |||||||
| wikipedia.org | |||||||
| yahoo.com | |||||||
| youtu.be | |||||||
| zoom.us | |||||||
| Domains scraped | 68 | 12 | 6 | 73 | 67 | 10 | 13 |
✓ The provider scrapes the domain. ✅ The provider scrapes the domain and returns structured JSON. ✕ The provider does not scrape the domain. Of the providers in this table, only Bright Data, Nimble and Zyte ran in the web unblocker benchmark; for the other providers the table shows JSON support only.
Web data providers markdown output performance
Markdown extraction success rate and completion time from our web unblocker benchmark, and the output formats each provider returns.
| Provider | Markdown | HTML |
|---|---|---|
Success rate by page type and concurrency
Product pages and search pages behave differently on the same domain.
Frequently asked questions
Methodology
The index re-uses the raw results of every web data benchmark we published. A benchmark counts if it tests web data providers and records the outcome of each request. A provider's index score is the average of its success rates across the benchmarks it ran in. Every benchmark carries equal weight, so a 1,400-request study counts as much as a 260,000-request one. Success means the response returned the target page with the content we asked for. We check that against bot pages and empty shells, so a 200 response with a challenge page counts as a failure. If a provider was never tested in a benchmark, we show a dash instead of a zero.
Targets come from the Tranco list, which ranks domains by averaging several traffic rankings. We remove dead, adult, gambling and malicious hosts before testing. Domains are grouped into categories: e-commerce, social media, search engines, video and streaming, travel, reviews, food delivery, app stores, real estate, job postings, news and publishing, developer and SaaS, finance, and government and education. Domain coverage and structured JSON support are checked on the top 100. Cost is per 1,000 successful pages, not per 1,000 requests sent. We take the cheapest plan a provider offers at that volume and divide by the success rate we measured. Completion time is how long one request takes from start to finish. Metadata fields are the parsed fields a provider returns when it answers with JSON. AI ready is the share of responses that came back as clean markdown. The anti-bot results come from two of our runs, the unblocker benchmark across 10,000 domains and the e-commerce benchmark across the top 100 domains, with each domain labelled by the vendor guarding it. Whiskers are 95% confidence intervals.
Explore Web Data Scraping Benchmarks
Ethical & Compliant Web Data Benchmark
As enterprises scale their web data operations, compliance, data, and risk executives increasingly evaluate the associated ethical, reputational, and legal risks. We benchmarked 5 leading web data collection services across 3 dimensions and tested each service with more than 20 potentially unethical scenarios. Our work helps you assess the ethical standing of your data collection…
Top 5 Job Posting Scraper APIs Compared
We benchmarked 5 leading web scraping providers across 5 major job platforms by running 12,500 requests in total, then measured each provider’s success rate, completion time, and metadata output. You can read benchmark methodology section for more details on the testing process = supported, returns HTML = supported, returns structured data = no data returned…
Top 5 Home Depot Scrapers Benchmarked & Compared
We benchmarked five web data providers on Home Depot, each fetching the same 50 product and search pages at 5 concurrent requests, for a total of 250 requests. You can read more about our benchmark methodology. Bright Data offers a dedicated scraper API for Home Depot, Apify provides a general e-commerce actor, and SerpApi also…
Best ScrapeBox Alternatives
ScrapeBox is a Windows and macOS desktop application used for SEO tasks such as search engine scraping, keyword harvesting, link building, comment posting, and backlink checking. However, it is a desktop GUI tool, not an API, and the cost is higher once premium plugins, proxies, and a CAPTCHA service are added. So alternatives are grouped…
eBay Scraping: Top 6 Providers Compared
We benchmarked eBay across 4 web scraping providers, totaling 1,400 requests over search and product pages of 7 eBay country sites, measuring both success rate and end-to-end completion time. You can read more about our benchmark methodology. Bright Data, Apify and SerpApi were the three providers that returned parsed JSON. Field counts came to 59…
Top Firecrawl Alternatives: Features & Pricing
Firecrawl is an AI-native web scraping and crawling API that converts dynamic web pages into LLM-ready Markdown and structured data. We compared its alternatives across benchmarks and features. We left out JavaScript rendering, search endpoints, structured JSON output and MCP server support. Nearly every product here offers all four, so they no longer separate one…
Top 7 Video Scrapers: Tested & Ranked
We tested the top 7 video scraping providers to see how they handle video metadata on the top video platform, totaling 6,000 requests, and measured their success rate, response time, and metadata fields. To see how we calculated these metrics, read video scraping benchmark methodology. Different providers return different amounts of metadata for the same…
Top 6 LLM Scrapers: ChatGPT, Perplexity & Gemini
We benchmarked how the top LLM scraper providers, including Bright Data, Oxylabs, and Apify, perform at extracting outputs from LLM platforms such as ChatGPT, Gemini, Perplexity, and Google AI Mode. To ensure reliable results, we ran 1,000 tests per provider, repeating each prompt 10 times for consistency. The top-performing provider is detailed below. Providers missing…
How to Bypass CAPTCHA (reCAPTCHA & hCaptcha)
Modern CAPTCHA and human-verification systems use a mix of challenge-response tests, browser signals, server-side token validation, and adaptive challenges. Attempting to bypass CAPTCHA on third-party websites can violate the terms of service or trigger account or IP blocks. The better approach is to use official APIs, reduce request rates, or implement a modern bot-management solution…
The Most Common Web Scraping Challenges
Web scraping has become more difficult in recent years. Since 2025, AI-related scraping has raised significant legal concerns. Platforms and infrastructure providers have adopted new methods to control AI crawlers and manage data collection. There are many technical challenges that web scrapers face due to the barriers set by data owners or website owners to…