Web Data Scraping Benchmarks
Provider success rates, response times and metadata coverage across our web scraping benchmarks.
Leaderboard
Filter by tool type and must-have requirements. Sort any column.
# | Model | Index | E-commerce | Social media | Search engines | Video & streaming | Travel | Reviews | Food delivery | App stores | Real estate | Job postings | News & publishing | Developer & SaaS | Finance | Government & education | Other |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Bright Data All tool types | 86.7 | 81.4 | 70.4 | 60.6 | 98.5 | 98.9 | 78 | 64.1 | 99.8 | 88.5 | 90.6 | 97.3 | 96.3 | 91.5 | 91.8 | 96.8 |
| 2 | Oxylabs All tool types | 77.2 | 19.8 | 70.4 | 0 | 99.7 | 89.1 | 56.2 | 91.7 | 99.6 | 71.1 | 77.5 | - | - | - | - | - |
| 3 | Apify All tool types | 75.4 | 18 | - | - | 99.6 | 99.2 | 98.6 | 44.6 | - | 94 | - | - | - | - | - | - |
| 4 | Nimble All tool types | 74 | 71.7 | 64 | 0 | 97.7 | 19.3 | 52.4 | 68 | 99.7 | 86 | 69.6 | 99.6 | 99.5 | - | 97 | 96.9 |
| 5 | Zyte All tool types | 72.5 | 25.5 | - | 6 | 99.6 | 96.3 | 65 | 75.9 | 99.5 | 65.6 | 57.8 | 97.6 | 98.6 | - | 96.7 | 97.8 |
| 6 | Decodo All tool types | 69.5 | 39.6 | 70.1 | 0 | 99.5 | 83.2 | 35.9 | 91.3 | 99.7 | 74.8 | 77.4 | - | - | - | - | - |
Cost vs. performance
Web scraper cost across our 100 e-commerce domain benchmark, and web unblocker cost across 10,000 domains. Volume-adjusted per successful request.
Avg is the effective per-1k rate we paid in our benchmark, volume-adjusted with each provider's cheapest available plan. Min is the cheapest tier each provider offers. Max is the most expensive tier.
Provider profile
Price per successful page, completion time, metadata depth, domains returning structured JSON and AI readiness for each provider.
# | Model | Domains with JSON | Avg metadata fields | $ per 1k | Completion time (s) | AI ready |
|---|---|---|---|---|---|---|
| 1 | Bright Data All tool types | 24 | 28 | 1.5 | 18 | 79 |
| 2 | Oxylabs All tool types | 12 | 4 | 2 | 14.2 | 71 |
| 3 | Apify All tool types | 10 | 42 | 13.7 | 26.7 | 55 |
| 4 | Nimble All tool types | 6 | 0 | 1 | 12.6 | 68 |
| 5 | Zyte All tool types | 0 | 0 | 3 | 13.6 | 60 |
| 6 | Decodo All tool types | 6 | 0 | 2.4 | 16.5 | 64 |
| 7 | SerpApi All tool types | 13 | 28 | - | 1.1 | - |
Behaviour under load
Success rate tracked at 1, 100, 500 and 5,000 parallel requests.
Anti-bot landscape
Which vendor guards your targets predicts your success rate.
Web unblocks across the top 10k domains
Scraping APIs across the top 10k domains
JSON coverage
Each provider was tested against the Tranco top 100 domains for whether it returns structured JSON rather than raw markup. A double check mark means the provider returned parsed structured fields for that domain; a cross means it did not. Domains are sorted by coverage, so the ones supported by the most providers appear at the top.
| Domain | |||||||
|---|---|---|---|---|---|---|---|
| google.com | |||||||
| chatgpt.com | |||||||
| facebook.com | |||||||
| youtube.com | |||||||
| amazon.com | |||||||
| bing.com | |||||||
| tiktok.com | |||||||
| googlevideo.com | |||||||
| microsoft.com | |||||||
| apple.com | |||||||
| instagram.com | |||||||
| fbcdn.net | |||||||
| twitter.com | |||||||
| office.com | |||||||
| github.com | |||||||
| wikipedia.org | |||||||
| youtu.be | |||||||
| yahoo.com | |||||||
| zoom.us | |||||||
| baidu.com | |||||||
| Domains with JSON | 24 | 12 | 6 | 6 | 0 | 10 | 13 |
Across the Tranco top 100 domains, Bright Data returns structured JSON for 24, SerpApi for 13, Oxylabs for 12, Apify for 10, and Decodo and Nimble for 6 each - Zyte returned none in this sample.
Success rate by page type and concurrency
Product pages and search pages behave differently on the same domain.
Start here
No provider wins every column. Choose the constraint that actually binds you.
Methodology
The index re-uses the raw results of every web data benchmark we published. A benchmark counts if it tests web data providers and records the outcome of each request. A provider's index score is the average of its success rates across the benchmarks it ran in. Every benchmark carries equal weight, so a 1,400-request study counts as much as a 260,000-request one. Success means the response returned the target page with the content we asked for. We check that against bot pages and empty shells, so a 200 response with a challenge page counts as a failure. If a provider was never tested in a benchmark, we show a dash instead of a zero.
Targets come from the Tranco list, which ranks domains by averaging several traffic rankings. We remove dead, adult, gambling and malicious hosts before testing. Domains are grouped into categories: e-commerce, social media, search engines, video and streaming, travel, reviews, food delivery, app stores, real estate, job postings, news and publishing, developer and SaaS, finance, government and education, and other. Structured JSON coverage is checked on the top 100. Cost is per 1,000 successful pages, not per 1,000 requests sent. We take the cheapest plan a provider offers at that volume and divide by the success rate we measured. Completion time is how long one request takes from start to finish. Metadata fields are the parsed fields a provider returns when it answers with JSON. AI ready is the share of responses that came back as clean markdown. The anti-bot results come from two of our runs, the unblocker benchmark across 10,000 domains and the e-commerce benchmark across the top 100 domains, with each domain labelled by the vendor guarding it. Whiskers are 95% confidence intervals.
Explore Web Data Scraping Benchmarks
Etsy Scrapers: Benchmarked Top 4 APIs
We benchmarked 4 web scraping providers on etsy.com. The results were filtered from our e-commerce scraping benchmark covering 100 e-commerce domains. A total of 200 requests were sent to Etsy at concurrency 5. You can read more about our Etsy scraping benchmark methodology. You can scrape either product or search pages and get full product…
Top 4 Google Play Scraping Providers Compared
We benchmarked four web scraping providers across Google Play product page URLs, sending 4,000 requests in total. For each request, we measured how reliably the provider returned data, how long it took from submission to final response, and how many metadata fields the response contained. Vendors are ranked by the number of structured metadata fields…
Top 7 Video Scrapers: Tested & Ranked
We tested the top 7 video scraping providers to see how they handle video metadata on the top video platform, totaling 6,000 requests, and measured their success rate, response time, and metadata fields. To see how we calculated these metrics, read video scraping benchmark methodology. Different providers return different amounts of metadata for the same…
Web Crawler Benchmark to Feed Websites to AI
We benchmarked four crawl APIs across three domains of varying difficulty at three max depth levels (5, 10, 20) with a 1,000-page limit, measuring crawl coverage, execution time, link discovery, markdown link quality, and title extraction accuracy. If you aim to: You can read our benchmark methodology. Firecrawl consistently crawled around 100 pages on theregister.com…
Top 6 LLM Scrapers: ChatGPT, Perplexity & Gemini
We benchmarked how the top LLM scraper providers, including Bright Data, Oxylabs, and Apify, perform at extracting outputs from LLM platforms such as ChatGPT, Gemini, Perplexity, and Google AI Mode. To ensure reliable results, we ran 1,000 tests per provider, repeating each prompt 10 times for consistency. The top-performing provider is detailed below. Providers missing…
Best Twitter (X) Scrapers: Benchmarked
We benchmarked the top Twitter (X) scrapers across 1000 URLs, for a total of 5000 requests. To help you choose the right tool for your Twitter scraping projects, we have categorized the top performers below. Since all providers reached 100% success rate, we compared their completion time. See our benchmark methodology for more details. When…
Best TikTok Scrapers: Scrape Video & Profile Data
A TikTok scraper collects public data from TikTok, including video metadata, profile details, engagement metrics, and comments, without using TikTok’s official API. We tested Bright Data, Apify, and Decodo by running 500 unique TikTok video URLs per provider. We measured two dimensions: validation success rate and the breadth of available metadata fields. See our methodology…
5 Best Scraping Browsers
Scraping browsers handle the unblocking infrastructure, enabling users to interact with websites programmatically and extract data easily. We benchmarked the top scraping browsers on sites with login walls, infinite scroll, and strict anti-bot rules. We updated this guide to include the latest anti-bot evasion techniques (TLS 1.3 fingerprinting) and updated pricing models for Bright Data…
How to Bypass CAPTCHA (reCAPTCHA & hCaptcha)
Modern CAPTCHA and human-verification systems use a mix of challenge-response tests, browser signals, server-side token validation, and adaptive challenges. Attempting to bypass CAPTCHA on third-party websites can violate the terms of service or trigger account or IP blocks. The better approach is to use official APIs, reduce request rates, or implement a modern bot-management solution…
The Most Common Web Scraping Challenges
Web scraping has become more difficult in recent years. Since 2025, AI-related scraping has raised significant legal concerns. Platforms and infrastructure providers have adopted new methods to control AI crawlers and manage data collection. There are many technical challenges that web scrapers face due to the barriers set by data owners or website owners to…