Plain HTML scraping typically runs $0.20 to $1 per 1,000 requests, JavaScript rendering pushes that to $1 to $5, and stealth or residential proxy configurations can hit $5 to $15 or more per 1,000 calls, with exact multipliers set by each vendor. The worked examples and calculator steps below turn that rule into a number you can put in a budget.
TL;DR:
- Costs for scraping vary widely based on configuration, from about $2-10 for small HTML requests to over $5,000 for large-scale stealth or residential proxies monthly.
- Per-request expenses depend heavily on multipliers such as JavaScript rendering, proxy type, and anti-bot measures, which can increase costs 5 to 75 times.
- Engineering maintenance and infrastructure, including proxies and data storage, often surpass the initial request fees at scale, making DIY less cost-effective long-term.
- Using spending caps, validation, and targeted data extraction can significantly reduce unexpected overages and control budget unpredictability.
- Transparent billing models and explicit failure responses help prevent surprise costs caused by retries, site errors, or misconfigured jobs.
Table of Contents
- Quick cost ranges and sample real-world numbers by configuration and volume
- What drives per-request cost: multipliers, target-site factors, and hidden billing rules
- Common pricing models: per-request, per-successful-request, per-GB, subscription and credits explained
- Total cost of ownership: engineering time, maintenance tax, proxies, compute, and QA
- How to estimate your bill: a step-by-step calculator and worked examples
- Practical cost-control tactics for developers and small data teams
- How transparent billing and failure modes reduce unexpected scraping costs
- Honest trade-offs between DIY, scraping APIs, and managed services
- Closing brief: fixing unpredictable scraping bills
- Sources
- FAQ
Quick cost ranges and sample real-world numbers by configuration and volume
Vendors advertise credit pools, not dollars, and that gap is where budgets get blown. A plan that claims 100,000 credits sounds generous until you learn that JavaScript rendering costs 5 credits per call and stealth mode costs up to 75, per credit multiplier data from HProxy. Translated into requests, that same 100,000-credit plan delivers 100,000 plain HTML pages, roughly 20,000 rendered pages, or as few as 1,333 pages on the highest stealth tier.
Here is what that looks like in practice across common configurations and monthly volumes:
- Plain HTML, 10,000 requests: roughly $2 to $10 per month at typical per-1k rates.
- JS-rendered, 100,000 requests: roughly $100 to $500 per month.
- Residential/stealth, 1,000,000 requests: roughly $5,000 to $15,000 or more per month, depending on target difficulty.
- Premium proxy blends: often sit between the JS-rendered and stealth tiers, since many targets only need premium IPs, not full anti-bot bypass.
A single request can cost anywhere from 1 to 75 credits depending on rendering, proxy type, and anti-bot settings, according to credit tables published by HProxy.
What drives per-request cost: multipliers, target-site factors, and hidden billing rules
Cost per request is a function of configuration, not raw page count. The multipliers stack, and each one you enable multiplies your effective credit burn against the plan you're paying for.
- JavaScript rendering typically multiplies base cost by around 5x, since it requires a headless browser instance instead of a raw HTTP fetch.
- Residential or premium proxies add another 10x-class multiplier on top of rendering, per HProxy's mapped weight tables.
- Anti-bot bypass or stealth mode can push a single call to the top of the range, up to 75 credits in some vendor configurations.
- Domain-specific weights exist too. Search engines and social networks are frequently priced higher than an ordinary retail or news site because they actively fight automated access.
Failed responses matter just as much as successful ones. Many vendors bill for 404s, 410s, and timeouts the same as a successful fetch, which means a stale URL list quietly burns credits on pages that no longer exist. Payload size, CAPTCHA solving, and geo-targeted proxy pools all add their own line items on top of the base multiplier, and none of them show up until the invoice does.
Pro Tip: *Before running a full job, fetch a random 20-URL sample from your target list and check the failure rate.
Common pricing models: per-request, per-successful-request, per-GB, subscription and credits explained
Vendors describe their pricing in different units, and the unit you're billed on changes how predictable your costs are.
- Per-request charges for every call attempted, successful or not, which is simple to reason about but punishes retry-heavy jobs and dead URLs.
- Per-successful-request charges only when data comes back clean, which raises the per-unit price but caps the damage from retries and failures, a trade some vendors describe explicitly in their pricing pages.
- Per-GB bills on data transferred, which fits large-asset jobs like image or PDF harvesting but can spike unpredictably on pages with heavy embedded media.
- Flat subscription bundles a fixed credit allotment into a monthly fee, which is easiest to budget around but wastes money if you consistently use less than the tier provides.
- Credit-based metering is the most common hybrid: a subscription base plus pay-as-you-go overage, where the credit-to-request math depends entirely on the multipliers covered above.
Small metadata extraction jobs usually fit per-request or per-successful-request billing best. Full-page HTML capture at volume tends to favor per-GB or a higher subscription tier. Whatever model you pick, check the vendor's overage and pause policy: some vendors halt service automatically at a spending cap, others bill the overage and invoice you later, and that difference alone can be the gap between a predictable bill and a surprise one.
Total cost of ownership: engineering time, maintenance tax, proxies, compute, and QA
The per-request price is rarely the real cost. Maintenance alone can consume 20% to 40% of an engineer's working time once a DIY scraper runs against a handful of changing targets, according to Tendem.ai's pricing analysis. At a fully loaded engineering salary, that maintenance tax alone can outweigh the direct cost of most vendor invoices.
Beyond salary time, a self-built pipeline carries recurring line items that never appear on a per-request quote:
- Proxy bandwidth for residential or premium pools, billed separately from the scraping logic itself.
- Headless browser compute, since rendering JavaScript at volume requires provisioned infrastructure, not a lightweight function call.
- Storage and monitoring for the extracted data plus the logging needed to catch silent failures before they compound.
- QA cycles to catch selector drift when a target site redesigns its markup, which happens without warning.
At real scale, the numbers get large fast: industry analyses put the fully loaded cost of collecting 10 million protected pages at $40,000 to $80,000 or more once proxies, retries, and engineering time are counted alongside raw request fees. That crossover point, where a managed API becomes cheaper than DIY once maintenance is priced in, usually arrives well before most teams expect it.
How to estimate your bill: a step-by-step calculator and worked examples
Run these steps against your own target list before committing to a plan:
- Count your targets and separate them into distinct site groups, since pricing varies by domain.
- Classify each group's complexity: 1x for plain HTML, 5x for JS rendering, 10x for premium proxies, 25x or more for stealth or anti-bot bypass.
- Estimate a realistic success rate from a small pilot sample, not the vendor's advertised uptime.
- Divide your target count by that success rate to get the true number of billed attempts, including retries.
- Multiply attempts by the per-request rate for your complexity tier to get raw request cost.
- Add engineering allocation (even 10% to 20% of one engineer's time for monitoring) if you're running the pipeline yourself.
Drop the success rate to 60% or add stealth mode, and that same job can double or triple.
Practical cost-control tactics for developers and small data teams
Most cost overruns come from re-fetching data you already have or retrying jobs blindly. A few habits close most of that gap:
- Use HTTP caching and conditional GETs so unchanged pages don't count against your request budget twice.
- Scrape deltas only when monitoring a target for changes, rather than re-pulling the full page on every run.
- Set spending caps and workspace budgets so a bug or a misconfigured crawl can't turn into a five-figure invoice overnight.
- Extract only the fields you need through schema-based extraction instead of pulling full HTML and parsing client-side, which can cut bandwidth and processing costs by 40% to 80% on typical jobs.
- Validate URL lists before a full run to catch dead links early, and pilot on 3 to 5 targets before scaling to the full list, a tactic covered in more detail in this guide to piloting scraping jobs.
Pro Tip: Run your pilot against the worst-case target in your list, not the easiest one. That number is the one your budget needs to survive.
Caching tactics and header-based fetch strategies are covered in more depth in this piece on HTTP caching for scraping pipelines.
How transparent billing and failure modes reduce unexpected scraping costs
This service sets spending caps at the workspace level, so a bug or a runaway crawl stops at a dollar figure you choose, not one you discover on an invoice. Every call returns a typed response, meaning failures like rate limits, CAPTCHA walls, or DNS errors are labeled explicitly instead of masquerading as generic errors that trigger blind retries.
- Spending caps stop a misconfigured job before it burns through a month's budget in an afternoon.
- Typed failure modes let your code branch on the actual error (rate limit versus dead domain) instead of retrying everything the same way.
- Schema-guided, LLM-powered extraction returns only the structured fields you asked for, which shrinks payload size and downstream processing cost.
The full primitive set and response shapes are documented at Gyrence's developer docs, and the console is open for testing against your own target list.
Honest trade-offs between DIY, scraping APIs, and managed services

DIY scraping makes sense when your volume is small, your targets are stable, and you have engineering time to spare. Beyond a few thousand requests a day against sites that change their markup, that math flips fast: the maintenance tax outpaces the infrastructure savings.
A scraping API fits most small to mid-size teams that want predictable per-request math without owning proxy rotation or headless browser fleets. Full managed services fit teams with zero engineering bandwidth to spare and a tolerance for paying a premium for someone else's QA. Decide fast with three questions: how soon do you need data flowing, how predictable does your budget need to be, and what happens to your pipeline the day a target redesigns its page.
— Glen
Closing brief: fixing unpredictable scraping bills
Gyrence was built around the exact problem this article walks through: vendors that hide multipliers until the invoice arrives. Every plan comes with workspace spending caps, so your monthly ceiling is a number you set, not one you discover.
- Spending caps and predictable credit math replace guesswork with a number you control before the job runs.
- Typed failure modes surface CAPTCHA, rate limit, and dead-domain errors explicitly, so retries are a choice, not a default.
- Schema-guided extraction returns only the fields you need, cutting payload size and downstream processing cost.
Compare tiers on the Gyrence pricing page, or check WebDoppler if you need change monitoring with webhook alerts layered on top of your existing scraping jobs.
Sources
Pricing figures and multiplier tables in this article draw on HProxy's credit multiplier breakdown, Tendem.ai's cost pricing guide, Titannet's scale-cost analysis, and AWS Lambda's official pricing page for a serverless request-cost baseline. For a developer-oriented model of clear API pricing presentation, see Chaingateway's pricing page.
- Scraping API Credits: 1 Request Can Cost 75 | HProxy
- Web Scraping Cost at Scale: How to Reduce Large-Scale Data Collection Costs | Titannet
FAQ
How much does it cost to scrape data?
Costs typically range from $0.20 to $1 per 1,000 requests for plain HTML, $1 to $5 per 1,000 for JavaScript-rendered pages, and $5 to $15 or more per 1,000 for stealth or residential proxy configurations, based on HProxy's published multiplier tables. Your actual bill depends on which configuration your targets require and your real success rate, not the vendor's advertised one.
How much does scraping cost at scale?
At large volumes against protected targets, fully loaded costs including proxies, retries, and engineering time can run $40,000 to $80,000 or more for around 10 million pages, according to Titannet's scale analysis. Engineering maintenance alone can consume 20% to 40% of a developer's time once a DIY pipeline runs at that scale, per Tendem.ai's research.
Is AI scraping illegal?
Whether AI-assisted scraping is legal depends on factors like the target site's terms of service, the data's copyright status, and whether the content is publicly accessible, and no single rule covers every case. A closer look at the legal axes that actually matter for developers is available in this legality primer.
Is web scraping illegal?
Scraping publicly accessible data is generally not illegal on its own, but violating a site's terms of service, bypassing access controls, or scraping copyrighted or personal data can create legal exposure depending on the jurisdiction and target. For a breakdown of the specific factors that determine legality, see this guide to web scraping legality.

