← Back to blog

Fix Scraping API Errors for Developers: 3 Checks and Typed Failures

October 3, 2026
Fix Scraping API Errors for Developers: 3 Checks and Typed Failures

When a scraping API throws an error, check three things immediately: the HTTP status code, the response body for a machine-readable error field, and whether a Retry-After header is present. Classify the result fast: 429, 408, and 5xx are retryable with backoff; 401, 403, and most other 4xx codes mean something in your request needs fixing first. Capture logs and the request ID before touching your client code.


TL;DR:

  • Retry responses should only be attempted for 429, 408, and 5xx errors, with proper backoff and jitter, and never for client (4xx) errors that indicate request flaws.
  • Always respect the server's provided Retry-After header for rate limits and back off accordingly, especially during 429 and 503 responses, rather than guessing delays.
  • Most permission issues (401/403) stem from token expiration, scope misconfiguration, or IP blocks, which should be fixed before retrying, not from persistent failures.
  • Reproduce errors with minimal requests and capture full request and response data, including headers and request IDs, to diagnose and escalate effectively.
  • Implement modest retry limits (three to five attempts), use idempotency tokens, and log each retry to avoid overwhelming the server or incurring unexpected costs.

Gyrence
Make Scraping Failures Actionable
Gyrence returns typed failure responses and spending caps, helping developers reason about web data requests without surprise scraping bills.
Explore Gyrence

Table of Contents

Common Scraping API Errors at a Glance

Most scraping failures fall into a handful of repeatable categories. The fix path depends entirely on which bucket the error lands in, and the status code is your fastest signal.

Client errors (4xx) usually mean something about the request itself is wrong, not the server:

  • 400 Bad Request: malformed query string, invalid JSON body, or missing required parameter. Fix the payload, don't retry blindly.
  • 401 Unauthorized: missing or expired credentials. Refresh the token before anything else.
  • 403 Forbidden: valid credentials but insufficient permission, or a bot-detection block. Check scopes first, then check for WAF signals.
  • 404 Not Found: the target URL or resource doesn't exist, or it moved. Confirm the URL hasn't changed.
  • 408 Request Timeout: the server gave up waiting on a slow client request. Usually safe to retry once.
  • 422 Unprocessable Entity: the request was well-formed but semantically invalid, such as an extraction schema that doesn't match the page. Fix the schema, don't retry.

Rate limiting gets its own category. A 429 Too Many Requests response often arrives with a Retry-After header specifying a wait time in seconds, sometimes ranging from a few seconds to an hour depending on the service. Respect that value exactly rather than guessing your own delay.

Server errors (5xx) point at the provider's infrastructure, not your request:

  • 500 Internal Server Error: generic failure on the provider's side. Retry with backoff.
  • 502 Bad Gateway / 504 Gateway Timeout: an upstream dependency failed or timed out. Retry, but escalate if it persists past a few attempts.
  • 503 Service Unavailable: often paired with Retry-After during maintenance or overload. Honor the header.

Two edge cases trip up developers more than they should. A 431 Request Header Fields Too Large usually means an oversized cookie or auth token got attached to the request, often from a proxy misconfiguration. A 511 Network Authentication Required signals a captive portal sitting between your client and the internet, not an API problem at all, and no amount of retrying the scraping endpoint will fix it.

A Compact Diagnostics Checklist: Reproduce, Capture, Classify

Before changing a single line of client code, turn the error into something you can reproduce on demand.

  1. Reproduce with a minimal request. Strip the call down to the smallest payload that still triggers the error, using curl or Postman instead of your full application stack.
  2. Capture the full exchange. Save the request headers, response headers, status code, and raw body, not just the error message your client library surfaces.
  3. Check for a Retry-After header and any rate-limit headers (X-RateLimit-Remaining, X-RateLimit-Reset), along with the request ID and server timestamp.
  4. Isolate the layer. Determine whether the failure is network-level (DNS, TLS handshake, a captive portal), API-level (status code from the provider), or client-level (a bug in your parsing or retry logic).
  5. Assemble a sanitized sample. Strip credentials and personal data from the payload, then attach it with timestamps and the request ID to any support ticket.

Pro Tip: Keep a small curl script per endpoint in version control. When an error shows up in production, you can reproduce it in under a minute instead of reconstructing the request from memory.

Retries and Backoff: Safe Retry Policies and Idempotency Rules

Not every error deserves a retry. Only transient failures, 408, 429, and the 5xx family, should trigger automatic retry logic. Everything else, including 400, 401, 403, and 404, needs a code or credential fix, not another attempt at the same broken request.

  • Use truncated exponential backoff with randomized jitter so retries don't cluster and hammer the server at the same intervals.
  • Cap retry attempts, typically three to five, and log each attempt with its delay and outcome.
  • Make write or extraction-triggering operations idempotent, or attach an idempotency token, so a retried request can't duplicate side effects.
  • Treat every retry as an event worth counting, not a silent background operation.

One documented vendor pattern: truncated exponential backoff with jitter, retrying only on 408, 429, and 5xx status codes, is the retry configuration recommended in Google Cloud's Vertex AI SDK documentation. The same principle applies to most scraping and extraction APIs: bound the attempts, randomize the delay, and never retry a request that failed for a reason retrying can't fix. Gyrence's retry guide walks through the implementation details for extraction workloads specifically.

Handling Rate Limits (429) and Quota Management

Persistent 429s usually mean your throughput assumptions don't match the provider's actual limits, not that the provider is unreliable.

  • Always honor a server-provided Retry-After value over any client-side guess; MDN documents it as the authoritative wait signal for 429 and 503 responses.
  • When no Retry-After header is present, fall back to conservative exponential backoff with jitter rather than retrying at a fixed interval.
  • Implement token-bucket or leaky-bucket throttling client-side so you never approach the limit in the first place.
  • Track usage per API key or per session, not just per IP address. OWASP's bot management guidance recommends layering identity and session signals because IP-only limits are easy to work around and easy to trip by accident behind shared NAT.

Pro Tip: If you're running LLM-powered extraction against a scraping API, account for compute cost per call, not just call frequency. A handful of heavy extraction requests can trigger secondary throttling even when you're well under the documented rate limit, because the server is reacting to CPU and memory load, not request count alone.

Expose quota usage and spending limits to whoever is operating the pipeline. Failing fast when a cap is hit, with a clear typed error, beats discovering an unexpected bill after the fact.

Authentication and Permission Errors (401/403): Validation and Fixes

Most 401 and 403 incidents trace back to one of four things:

  • Authorization header format. Confirm the scheme (Bearer, Basic, or a custom header) matches what the API expects, and check that the token hasn't expired since your last refresh.
  • Scope and role checks. A valid token can still return 403 if it lacks permission for the specific resource or operation being requested.
  • Signature and clock skew. For HMAC-signed requests, verify the signature is computed over the exact payload the server receives, and check that your system clock hasn't drifted, since many signing schemes reject requests outside a narrow timestamp window.
  • WAF or IP-level blocks. A 403 that shows up only from certain networks or IP ranges often points to a bot-management rule rather than anything wrong with your credentials; understanding when to use different proxy types can help mitigate these IP-level blocks, as explained in this guide to the best proxies for web scraping.

Fixing these in order, from token format to permission scope to signature to network-level blocks, resolves the overwhelming majority of 401/403 cases without needing provider support.

Parsing, Structural Drift, and Anti-Bot Blocks

A scraping failure that isn't an HTTP error is often a parsing problem. The target page changed its structure, and your selectors or extraction schema no longer match.

  • Distinguish a hard parser exception (malformed HTML or JSON that breaks your parser outright) from a silent failure where parsing succeeds but returns an empty or wrong field, since the second kind doesn't throw an error and is easy to miss.
  • Use tolerant parsing libraries and validate extracted JSON against a schema before trusting it downstream.
  • Watch for CAPTCHAs and bot-detection challenges early. A page that normally returns content but starts returning a challenge page is a signal to pause and route to an authorized API or a human-verification path, not to retry harder.
  • Keep a fallback extraction path and track drift rate over time so a layout change gets caught before it silently corrupts your dataset.

Pro Tip: Log a hash of the page structure alongside your extracted data. A sudden shift in that hash is often the earliest sign of drift, well before your extraction logic actually breaks.

CORS and Browser-Side Failures: Debugging Tips for Front-End Testing

CORS errors are enforced entirely by the browser, and the error message it gives you is deliberately vague. The actual cause lives in headers you can only see by inspecting the exchange directly.

  • Open DevTools Network tab and look at the preflight OPTIONS request first. MDN's CORS documentation explains that the browser hides server-side detail from your JavaScript, so the preflight response is your real source of truth.
  • Confirm the server sends the correct Access-Control-Allow-Origin, Access-Control-Allow-Methods, and Access-Control-Allow-Headers values. A missing or mismatched header is the usual culprit.
  • Simplify the request where possible (avoid custom headers or non-standard content types) to dodge the preflight step entirely.
  • When you can't change the server, route requests through a server-side proxy for local development instead of fighting the browser's CORS policy.
  • Recognize an opaque response from a no-cors fetch call: status 0 with no readable body. That's expected browser behavior, not a server error worth debugging on the client side.

Observability: What to Log, Metrics to Monitor, and Alert Rules

An error you can't reproduce is an error you can't fix. Logging the right fields turns a one-off incident into a pattern you can act on.

  • Log the request ID, timestamp, full request and response headers, latency, status code, and a sanitized sample of the payload for every failed call.
  • Track error rate by endpoint, 429 frequency, mean time between failures, and retry counts as standing metrics, not just one-off debug output.
  • Set alerts on rising 429 or 5xx rates and on abnormal latency spikes, with automatic escalation after a defined number of consecutive failures.

A useful reference point: RFC 6585 notes that servers using status codes like 429 have discretion in how they identify a client and count requests, which means the same error can mean different things across providers. Logging enough context to tell those cases apart is what makes an alert actionable instead of just noisy.

Sanitize personal data out of any payload sample before it goes into long-term storage. Keep it long enough for triage, not longer. Gyrence's data repeatability guide covers drift monitoring and long-term observability patterns in more depth for teams running extraction at scale.

Scraping public data is not automatically illegal, but certain tactics raise real legal risk: bypassing CAPTCHAs or other technical access controls, misrepresenting who or what is making the request, or pulling data that's private or access-gated. Enforcement actions in the United States have invoked the CFAA and FTC authority in cases involving deceptive or unauthorized access. Document your intent, avoid deceptive attribution, and get explicit permission before touching protected or private data. Gyrence's legal guide for developers covers the four axes that typically decide these cases.

Gyrence Approach: Typed Failures, Spending Caps, and Primitives

Gyrence structures its API around five composable primitives, Search, Traverse, Fetch, Extract, and Map, and every call returns a typed, discriminated-union response. That includes the failure cases: instead of parsing an unstructured error message, an agent or script can read a typed field and decide programmatically whether to retry, fall back to a different primitive, or escalate to a human.

Five API primitives feeding typed failure decisions

That structure matters most under load. When a 429 or a drift-related extraction failure happens mid-run, a typed response tells the caller exactly what kind of failure it's looking at instead of forcing a guess from status code and prose alone. Spending caps apply the same honesty principle to billing: usage stops at a defined limit instead of accumulating silently. For implementation-level detail, see Scraping Retries Explained and documentation on decoupling fetching from failure handling.

Author Perspective: Repeatable Triage and Modest Retries Beat Heroic Debugging

The instinct under pressure is to retry harder and faster. That instinct is usually wrong. A measured retry policy with solid telemetry catches more real problems than three parallel retry loops firing blind. Predictable billing controls, like a hard spending cap, remove the temptation to brute-force a flaky endpoint and make the incident easier to reason about afterward. The practical rule holds up across almost every scraping failure: reproduce it small, log it completely, retry it modestly, and escalate it clearly.

— Glen

Try Gyrence: Typed Failure Responses, Predictable Billing, and Docs

Typed, discriminated-union responses mean a failed call tells you what happened instead of leaving you to parse a vague error string. Spending caps mean a bad retry loop or a sudden traffic spike stops at a number you set, not whatever the provider decides to bill. Gyrence

Try Gyrence: Typed Failure Responses, Predictable Billing, and Docs — overview diagram

If you're building extraction pipelines or agent tooling and want to see how structured failure handling behaves in practice, the pricing page covers the Free, Founders, Pay-As-You-Go, Standard, Growth, and Scale plans, and WebDoppler is worth a look if ongoing drift monitoring and webhook alerts fit your use case.

Sources

FAQ

What are some common API errors?

The most frequent ones are 400 (bad request), 401 (unauthorized), 403 (forbidden), 404 (not found), 408 (timeout), 429 (rate limited), and the 5xx server error family. Each points to a different fix: credentials for 401/403, request structure for 400/422, and backoff for 429 and 5xx.

Is AI scraping illegal?

AI-assisted scraping follows the same legal principles as any other scraping: it isn't automatically illegal, but bypassing access controls, ignoring robots.txt signals, or extracting private or gated data increases legal risk. Review Gyrence's legal guide for the factors that typically matter.

Is web scraping illegal?

Scraping publicly available data is generally lawful, but specific tactics, like bypassing CAPTCHAs, violating a site's terms in certain ways, or accessing non-public data, can expose a developer to legal claims including CFAA or FTC action in the United States. The safest path is documenting intent and seeking permission for anything outside clearly public data.

What does scraping an API mean?

Scraping an API means programmatically calling an API's endpoints to pull structured data, rather than parsing raw HTML from a web page. It still requires handling the same class of errors, authentication failures, rate limits, and server errors, that any HTTP client needs to manage.