10 Web Scraping APIs Compared: Output Formats, Rendering, Async Jobs, and Cost Per Usable Record

Justin Shin

The right web scraping API isn't the one with the boldest marketing slogan — it's the one whose output format, rendering method, and job model match what your pipeline actually needs, at a cost per valid record you can verify yourself rather than take on faith.

Most comparisons of scraping APIs lean on speed claims, "fastest" or "cheapest" labels, and pricing snapshots that go stale within months. None of that tells you whether a given API returns raw HTML, a JSON wrapper around HTML, or genuinely parsed fields — and that distinction determines how much post-processing work lands on your side. Below is a breakdown of ten providers by output contract, rendering and anti-bot approach, and job model, followed by a practical method for validating output and calculating what a usable record actually costs.


Table Of Contents

Why the Right Choice Depends on Output Contract, Not Slogans

Every provider on this list handles proxies, and most handle some mix of captchas and browser rendering. Where they genuinely differ is in what lands in the response body. Some return raw HTML you still have to parse. Some wrap HTML inside a JSON envelope, which looks structured but isn't parsed data. Others return actual named fields — price, title, availability — extracted server-side. Picking based on a headline claim instead of the response contract can create unexpected parsing work.


HTML, JSON Wrapper, or Parsed Fields: Read the Contract Carefully

According to Crawlbase's documentation, its current Crawling API can return HTML, a JSON wrapper, or Markdown, and a separate scraper library exists alongside an asynchronous Enterprise Crawler; the older standalone Scraper API is now marked legacy in the same docs. The important nuance: a JSON wrapper around a page is not the same as parsed product fields, so confirm which response type an endpoint actually returns before assuming you'll skip a parsing step.

ScrapingBee's documentation describes an HTML API with JavaScript rendering plus extraction rules built on CSS and XPath selectors, and browser scenarios for interactive pages — configuration choices that affect credit consumption. Zyte API's current documentation supports automatic extraction for product pages, articles, job postings, and general page content, sourcing data from either the raw HTTP response or rendered browser HTML depending on configuration; this replaces the older AutoExtract branding, and returned fields still need validation for your target pages.


JavaScript Rendering and Anti-Bot Handling

When required content exists only after browser execution, rendering may be needed. If an authorized API or the initial HTML already contains the data, a browser may add unnecessary cost. WebScrapingAPI's documentation splits this out explicitly, offering separate general HTML, browser-rendering, SERP, and e-commerce endpoints rather than one universal call — meaning the endpoint you pick determines both the response contract and the cost profile. ScraperAPI's documentation similarly separates HTML retrieval, an async API, and structured endpoints, with capabilities that must be checked for the selected endpoint. This comparison does not establish guaranteed access or detection outcomes on any target.


Async Jobs: An Accepted Request Isn't Finished Data

Several of these APIs run jobs asynchronously for heavier rendering or large-scale crawling. The pattern is consistent across providers: you submit a request, receive a job or run identifier, and then either poll an endpoint or listen for a webhook until the job reports completion. An HTTP 202 or a job ID means the request was accepted — not that data is ready. Apify's documentation describes running an Actor or task directly via REST and fetching results from a dataset once the run finishes; its SDK is optional, not required, contrary to the assumption that Apify demands installing a library to get data out. Crawlbase's Enterprise Crawler follows a similar accept-then-retrieve model for large async jobs. In practice, build in bounded retries and backoff per each provider's documented limits, and design for failure handling and deduplication, so a failed or delayed job does not silently become missing data.


10 Web Scraping APIs Compared

Provider Output Options Job Model Scope Notes
Crawlbase HTML, JSON wrapper, Markdown Sync + async Enterprise Crawler Separate scraper library; older Scraper API is legacy
ScraperAPI HTML, structured endpoints Sync + async API DataPipeline and SDKs; no universal target guarantee
ScrapingBee HTML with extraction rules Sync, with browser scenarios CSS/XPath extraction; rendering config affects credits
Apify JSON dataset items Actor/task runs, async SDK optional; output/billing vary per Actor
WebScrapingAPI HTML, JSON depending on endpoint Endpoint-specific; check async options Separate HTML/browser/SERP/commerce APIs
Scrapestack HTML Check selected endpoint Configurable JS rendering; no universal speed/cost claim
Bright Data SERP API Structured search results Check selected endpoint Search-engine specific, distinct from general scraping APIs
Decodo HTML, JSON, Markdown, CSV Sync + async templates Target templates determine available output
Shifter HTML, JSON Check selected endpoint Separate scraping and SERP API lines
Zyte API Automatic extraction fields, HTML Sync per request Extraction source varies by page type

Below is a short profile of each provider, in the same order as the table.

Crawlbase

Crawlbase's documented Crawling API returns HTML, a JSON wrapper, or Markdown, with a separate scraper library and an asynchronous Enterprise Crawler for larger jobs; its older Scraper API is now marked legacy in current docs. For targets outside its specialized parsers, expect raw HTML rather than parsed fields — the JSON wrapper is a delivery format, not automatic extraction. It sits alongside a broader proxy infrastructure, and pairs with related resources such as this site's Amazon scraping proxy guide if you're building around e-commerce targets.

ScraperAPI

ScraperAPI's docs describe HTML retrieval, an asynchronous API, structured endpoints for select targets, a DataPipeline product, and SDKs across several languages. Validate against your specific targets rather than assuming broad coverage from the product name.

ScrapingBee

ScrapingBee documents an HTML API with JavaScript rendering, CSS/XPath extraction rules, and configurable browser scenarios for interactive pages. Extraction rules reduce downstream parsing work for well-defined page structures, though more complex rendering configurations consume more credits per request according to its own docs.

Apify

Apify runs are triggered directly via REST against an Actor or task, with results retrieved from a dataset once the run completes — its SDK is optional, not a requirement for basic use. Because Actors are built by different developers, input parameters, output shape, and billing all vary by Actor; not every Actor is a scraper, and not every Actor returns JSON-only output, so check each Actor's own page before building around it.

WebScrapingAPI

WebScrapingAPI's current documentation separates general HTML, browser-rendering, SERP, and e-commerce endpoints rather than exposing one universal API. Choosing the correct endpoint for your target matters more than a single "fastest" label, since each endpoint carries its own response contract and credit cost.

Scrapestack

Scrapestack, documented on APILayer's official product page, offers REST-based HTML retrieval with JavaScript rendering and configurable options. Compare the billed configuration and usable output on your own sample; this review establishes no price or speed winner.

Bright Data SERP API

Bright Data positions itself as a web data infrastructure provider, and its SERP API documentation covers structured search results from Google, Bing, and other listed engines. This is a distinct product from a general-purpose web scraping API and is scoped specifically to search engine results pages. The documented feature set is the basis of this comparison; no independent performance ranking or permission to collect a particular dataset is implied.

Decodo

Decodo (formerly Smartproxy) documents a unified Web Scraping API built around target templates and rendering options, with output format — HTML, JSON, Markdown, or CSV — depending on the template selected. Pricing runs on credits tied to the options chosen per request. Older claims of fixed, narrowly limited social media or e-commerce targets are outdated relative to current docs. If you're already running proxy infrastructure, this pairs with broader guidance in this site's residential proxy comparison.

Shifter

Shifter's official site lists a Web Scraping API and a separate SERP API, both describing JavaScript rendering and proxy handling. These are vendor-described features on the provider's own page, not independent proof of throughput or reliability at scale — worth validating with your own sample before committing to a plan.

Zyte API

Zyte API's current documentation supports automatic extraction for product pages, articles, job postings, and general page content, pulling from either the raw HTTP response or rendered browser HTML depending on the request. This is the modern successor to the older AutoExtract branding. Validate field completeness and meaning against representative source pages before accepting automatically extracted records.


Validating Results and Calculating Real Cost Per Record

A 200 HTTP status indicates HTTP-level success at the responding service — it does not confirm the data is correct or complete. Before trusting a batch of results, check for:

Once you've filtered out invalid records, the only cost figure worth trusting is total billed cost divided by valid, unique, usable records — not the advertised per-request price. That formula should include retries, rendering surcharges, AI-assisted parsing, Actor compute, or export fees, since these vary by provider and are not part of a universal base rate. As a purely illustrative example: if 1,000 requests return 800 valid, usable records, your real cost per record is the total bill divided by 800, not by 1,000. Run that math on your own account statement for any provider before comparing plans.


A Practical Way to Test Before You Commit

This comparison did not run a controlled benchmark across all ten providers, and any speed or accuracy ranking you see elsewhere should be treated the same way unless it names its test set. A workable method: pick a small, authorized sample of URLs you're permitted to access, covering a static page, a JavaScript-heavy dynamic page, and a paginated listing. Run the same sample through each candidate API and compare completion time, field accuracy against manual inspection, and the cost-per-usable-record formula above. That gives you a result specific to your actual targets, which matters more than any generic vendor claim.


Public availability of a webpage is not automatic permission to scrape and reuse everything on it — that depends on the site's terms, applicable law, and the nature of the data involved, and none of that is settled by a scraping API's marketing copy. Where a target site offers its own official API and it meets your data needs, that's generally the more direct and stable route. Operationally, keep API keys and credentials server-side rather than exposed in client code, and minimize the fields and retention period for any personal or sensitive data you collect. None of this is a workaround for access controls, and the ability to use an API does not establish authorization for a target.


Bottom Line

Choose by matching output format, rendering needs, and job model to your actual pipeline — then confirm the real cost per validated record yourself, since that number, not a headline claim, is what determines whether a scraping API is worth the spend.

Crawlbase, ScraperAPI, ScrapingBee, Apify, WebScrapingAPI, Scrapestack, Bright Data's SERP API, Decodo, Shifter, and Zyte API each solve a slightly different piece of the scraping problem — some return raw HTML you'll parse yourself, some wrap it in JSON, others extract structured fields automatically, and Bright Data's product is scoped specifically to search results rather than general web pages. Read each provider's current documentation for the exact response contract, test against a small authorized sample of your own targets, and calculate cost against valid usable records rather than advertised request pricing before committing to a plan.

Featured Reviews

18,410 Reviews Analyzed
162 Reviews Analyzed
59 Reviews Analyzed
385 Reviews Analyzed
94 Reviews Analyzed
64,795 Reviews Analyzed
79,497 Reviews Analyzed
2 Reviews Analyzed
168 Reviews Analyzed

Related Posts