Default to HTTP requests. They're faster, cheaper, and easier to scale, and they handle the majority of structured data extraction without touching a browser engine. Reach for a headless browser only when the page genuinely requires JavaScript execution, DOM rendering, or multi-step interaction. The decision checklist below tells you which situation you're in before you write a line of code.
TL;DR:
- HTTP requests are ideal for static data loaded directly in the initial HTML or through lightweight network calls, saving cost and complexity.
- Use headless browsers only when content appears after JavaScript execution, user interactions, or behind complex client-side frameworks that lack accessible APIs.
- Replicate API calls directly when possible, as calling JSON endpoints reduces bandwidth use and improves reliability compared to full page rendering.
- Running headless browsers demands significant resources, introduces complex failure modes, and scales less predictably than simple HTTP requests.
- A disciplined workflow involves inspecting network traffic in DevTools first and escalating to headless browsers only when API endpoints are unavailable or dynamic rendering is essential.
Table of Contents
- Quick decision guide: HTTP or headless?
- How HTTP-based scraping works: fetch, XHR, and hidden endpoints
- How headless browsers work and why they cost more
- Trade-offs: throughput, reliability, and cost
- Workflow: from network inspection to headless escalation
- Keeping production scrapers reliable: service workers, proxies, retries
- Why "browser execution last" is the safer default
- A managed alternative to running headless infrastructure yourself
- FAQ
- Sources
Quick decision guide: HTTP or headless?
Check the page before you pick a tool. If data loads in the initial HTML or arrives through a visible JSON or XHR call in the network tab, HTTP fetch is enough. If content only appears after JavaScript runs, after scrolling, after a login flow, or behind a click, you need browser execution.
Signals that HTTP is sufficient:
- The page source (view-source) already contains the data you need.
- DevTools' Network tab shows a clean JSON or XML response powering the page.
- The site exposes a documented or guessable API endpoint.
Signals that demand headless:
- Content renders only after client-side JavaScript runs.
- The flow requires clicks, scrolls, or authenticated sessions with dynamic tokens.
- The target uses heavy client-side frameworks with no backing API you can call directly.
Rule of thumb: inspect the network tab first, prefer JSON endpoints always, and budget separately for headless infrastructure since it carries different cost math entirely.
How HTTP-based scraping works: fetch, XHR, and hidden endpoints
HTTP scraping means sending a direct request and reading the raw response, no rendering involved. The Fetch API is the standard way to do this in JavaScript environments, and most scraping libraries in other languages mirror the same request and response model.
A critical detail trips up a lot of developers: a fetch() promise only rejects on network failure. A 404 or 500 response still resolves successfully, so you have to check Response.ok or the status code yourself to catch HTTP errors after the call completes.
Finding the right endpoint takes a few steps:
- Open DevTools and load the page with the Network tab recording.
- Filter by XHR or Fetch to isolate the calls the frontend makes to populate data.
- Inspect the response body for the JSON payload you actually need.
- Replicate that request with the correct headers, cookies, or auth tokens.
Calling the JSON endpoint directly instead of rendering the whole page saves bandwidth and improves reliability, and you can cache responses at the HTTP layer with standard headers, something no browser-rendering pipeline gives you for free. HTTP fails when the data genuinely doesn't exist until a script runs client-side: at that point, no amount of clever header manipulation gets you there.
Pro Tip: Before writing any scraper code, spend five minutes in DevTools' Network tab. The endpoint you need is often already there, waiting to be called directly.
How headless browsers work and why they cost more
A headless browser runs a full browser engine, Blink in Chrome's case, without a visible window. It parses the DOM, executes JavaScript, manages storage, and runs service workers exactly as a normal browser session would. Headless Chrome shares the same rendering engine as headful Chrome, so it carries the same performance cost profile.
Tools like Playwright and Puppeteer follow a predictable lifecycle: launch the browser, create a context, open a page, navigate, wait for the target content, interact if needed, then extract.
- Launch and context creation alone consume meaningful CPU and memory before any page even loads.
- Each open page holds its own rendering state, which multiplies resource use under concurrency.
- Interactions like clicks or scrolls add timing dependencies that make scripts more fragile than a single HTTP call.
Modern headless Chrome runs the same code paths as headful Chrome, which means you should expect similar memory and CPU footprints when running at scale, not a lighter alternative. Headless unlocks capability, full JavaScript execution, screenshots, network interception, but that capability comes bundled with orchestration overhead. A browser that isn't configured correctly, missing wait conditions, unhandled dialogs, stale selectors, will fail quietly in production long before an HTTP client would.
Trade-offs: throughput, reliability, and cost
HTTP requests and headless sessions scale in fundamentally different ways. An HTTP client can fire thousands of concurrent requests limited mainly by your own rate limits and the target's tolerance. A headless browser is limited by available memory and CPU per instance, since each page carries its own rendering context.
- HTTP failures are usually clean: a status code, a timeout, a malformed body you can parse and log.
- Headless failures are murkier: a selector that silently returns nothing, a navigation that times out, a dialog that blocks the page.
- Proxy and session strategy differs too: HTTP scraping reuses connections and cookies cheaply, while headless sessions need full browser contexts rotated alongside proxies.
Debugging a stalled headless job often means replaying the session with a visible browser to see what actually happened, a step HTTP scraping rarely requires. Cost follows the same split: HTTP calls scale close to linearly with request volume, while headless infrastructure scales with concurrent browser instances, which is a steeper and less predictable curve as traffic grows. Scaling headless browsers is routinely underestimated because browser lifecycle and memory management make it more failure-prone than an equivalent HTTP pipeline.
Pro Tip: Log the exact failure type for every scrape attempt, timeout, selector miss, HTTP error, rather than a generic "failed." You can't fix what you can't categorize.

Workflow: from network inspection to headless escalation
A disciplined workflow keeps you from reaching for browser automation by default.
- Load the target page with DevTools open and watch the Network tab for JSON or XHR calls that supply the content.
- If you find one, replicate it with
fetch()or an HTTP client, matching headers likeAuthorization,User-Agent, and any required cookies. - Test the raw request outside your scraper first (curl or Postman) to confirm it returns the expected payload without a browser session.
- If no API call exists and the content only appears after JavaScript execution, escalate to a headless tool like Playwright, as detailed in this guide to scraping dynamic pages.
- Size your headless infrastructure for concurrent browser contexts, not request volume, since that's the real constraint.
This order matters because every step you skip toward headless costs you reliability and money later. A comparison of login-then-HTTP workflows shows how often a single headless step (logging in) can hand off to a pure HTTP session for the rest of the extraction.
Keeping production scrapers reliable: service workers, proxies, retries
Service workers are a common silent failure point. They can intercept and serve requests independently of the main page, which means some network traffic never shows up in your automation tool's route handlers. Playwright's network documentation recommends disabling service workers, or explicitly accounting for them, when expected requests seem to go missing during interception.
- Use
browserContext.route()to intercept and mock requests, but verify service workers aren't bypassing your routes entirely. - Rotate proxies per session rather than per request when using headless browsers, since a browser context is already a stateful unit.
- Reuse cookies and sessions across HTTP calls to cut both latency and the chance of tripping rate limits.
- Log structured failure signals (timeout, selector not found, HTTP status) so retries target the actual cause instead of blindly repeating the same call.
Pro Tip: Set retry backoff based on failure type: a 429 status needs a longer wait than a transient navigation timeout. For orchestration patterns at the agent level, a third-party breakdown of AI scraping agent tiers is worth a look if you're coordinating many scrapers under one cost ceiling.
Why "browser execution last" is the safer default
Teams default to headless because it feels safer, it handles anything a browser can render. That instinct creates technical debt: more infrastructure, murkier failures, and a cost curve that punishes scale. The disciplined move is checking for an HTTP path first every time, even when you're fairly sure you'll need a browser. Most of the time, you won't, and the few minutes spent in DevTools pay for themselves many times over in uptime and in the bill you don't get surprised by.
— Glen
A managed alternative to running headless infrastructure yourself
Running your own headless fleet means owning browser lifecycle management, proxy rotation, and failure classification indefinitely. A managed service can handle that layer: several composable primitives cover both HTTP-first and browser-execution cases through one API, with every response returned as a typed, failure-aware result instead of a silent empty page.
Extract features schema-guided JSON output built in so you're not stitching a separate extraction step onto your scraper. Billing runs on predictable, capped credit usage rather than a growing server bill, with plans from Free through Scale on the pricing page. For sites that change shape over time, WebDoppler adds drift monitoring with webhook alerts, so you find out about a broken selector before your downstream job does.
FAQ
Is web scraping illegal in the US?
Web scraping itself is not inherently illegal, but legality depends on what data you collect, how you access it, and the site's terms of service. Public, non-personal data generally carries lower legal risk than scraping behind authentication or collecting personal information, and reviewing ethical scraping practices around rate limits and PII is a reasonable starting point before building a scraper.
What are the disadvantages of using a headless browser?
Headless browsers consume significantly more memory and CPU than HTTP requests because they run a full rendering engine for every page. They also introduce more failure points, timing issues, selector changes, service worker interference, that make debugging and scaling harder than a comparable HTTP pipeline.
Is web scraping outdated?
No. Scraping remains a core method for structured data extraction, though the tools have shifted toward API-first approaches where possible and headless browsers where JavaScript rendering is unavoidable. Demand has grown further with AI agents and RAG pipelines that need fresh, structured web data on demand.
Is scraping with BeautifulSoup legal?
BeautifulSoup is a parsing library, not a legal classification, so its use is legal in the same sense any HTTP client is. The legal considerations are the same as with any scraping method: what data you access, how you access it, and whether you respect the target site's terms and applicable law.

