Treat rate limits as authoritative traffic signals: honor Retry-After, apply exponential backoff with jitter, and instrument monitoring before you scale. This sequence prevents bans and keeps extraction predictable. Skip any of the three and you trade a short-term speed gain for blocked IPs, wasted retries, and an API bill nobody can forecast.
TL;DR:
- Ignoring
Retry-After, exponential backoff, and monitoring leads to IP bans, unpredictable costs, and inefficient retries during rate limiting.- The primary rate limit signals include HTTP 429 responses with
Retry-Afterheaders and rate-limit headers likex-ratelimit-remainingandx-ratelimit-reset.- Different APIs use token buckets, fixed windows, or sliding windows to count requests, with additional secondary burst limits layered on top of primary ones.
- Implementing adaptive retries with jitter, bounded attempts, and retry quotas prevents storming and feature misuse, while testing logic offline avoids quota waste.
- Respecting crawl directives like
robots.txtandcrawl-delayreduces blocks, especially when using API endpoints with documented limits over public HTML scraping.
Table of Contents
- What rate limiting is and the key HTTP signals to watch for
- Server-side counting models and algorithms
- Client-side resilience: retries, backoff, and adaptive limiting
- Polite crawling: robots.txt, crawl-delay, and site directives
- Proxy and header strategies: rotation trade-offs and fingerprinting
- Testing and simulation: validating retry logic safely
- Monitoring, observability, and cost control
- Gyrence approach: predictable scraping with typed failures and spending caps
- Short practical rules we follow in production
- Make scraping predictable with a managed web data API
- FAQ
- Sources
What rate limiting is and the key HTTP signals to watch for
Servers throttle requests to protect capacity, stop abuse, and keep service fair across every client hitting the same endpoint. For a scraper, that throttling shows up as specific, readable signals, not a mystery wall.
The core signal is the HTTP 429 status code, which means the client sent too many requests in a given window. MDN's reference documents a 429 response that often carries a Retry-After header, for example Retry-After: 3600, telling the client exactly how many seconds to wait. RFC 6585 formalized this status code, recommending that servers include details explaining the condition and noting that 429 responses must never be cached.
Beyond the status line, well-behaved APIs expose rate-limit state directly in headers:
x-ratelimit-limit: the total requests allowed in the current window.x-ratelimit-remaining: how many requests are left before you hit the wall.x-ratelimit-reset: a timestamp marking when the window resets and your quota refills.Retry-After: a direct instruction, in seconds or a date, for how long to wait before your next attempt.
Reading these headers before every request, not just after a 429, is what separates a scraper that adapts from one that just keeps getting punished.
Server-side counting models and algorithms
Different platforms count requests differently, and the counting model determines how your client should behave under load.
- Token bucket: a bucket refills at a fixed rate and each request consumes a token; bursts are allowed as long as tokens exist, then throttling kicks in hard once the bucket is empty.
- Fixed window: requests are counted within a set block of time, like 1,000 per minute; a client that times a burst right at the window boundary can send double the intended rate across the two windows.
- Sliding window: a rolling calculation smooths out the boundary problem, counting requests over a continuously moving time frame rather than a fixed clock tick.
Many APIs layer secondary or burst limits on top of the primary one. GitHub's API, for example, imposes primary limits per hour alongside secondary limits that trigger on rapid concurrent requests, a pattern Microsoft's Dev Proxy documentation uses as its test case for simulating rate-limit behavior locally. These secondary limits exist specifically to stop short, aggressive bursts that a primary hourly count would otherwise miss.
Counting itself can key off several different identifiers:
- IP address: simple to implement, but Cloudflare notes that IP-based limiting produces false positives, since many legitimate users can share a NAT'd or corporate IP.
- API key or token: ties the limit to your account rather than your network, which is why authenticated endpoints often give higher, more predictable ceilings.
- JWT claim or user ID: lets a platform enforce per-user fairness even when many users share infrastructure.
- Resource path: some endpoints, like search or write operations, carry tighter limits than read-only ones because they cost more to serve.
Knowing which model and key a target API uses tells you whether spreading requests across time, across keys, or across both will actually help.
Client-side resilience: retries, backoff, and adaptive limiting
The baseline pattern is exponential backoff with jitter: each retry waits longer than the last, and randomness prevents every client from retrying at the exact same instant. AWS's retry guidance recommends a base delay of 1,000 milliseconds for throttling errors, a maximum backoff cap of 20 seconds, and full jitter, which randomizes the wait uniformly between zero and min(cap, base * 2^attempt).
A workable retry configuration looks like this:
- Base delay: start at a short delay for throttling responses; shorter bases suit transient network errors instead.
- Backoff cap: never let a single wait exceed a recommended maximum, typically around 20 seconds.
- Max attempts: bound retries to a limited number; an unbounded retry loop just delays the inevitable failure while burning quota.
- Jitter: always randomize within the capped window rather than retrying on a fixed clock, which avoids synchronized retry storms across concurrent workers.
Beyond basic backoff, a retry quota or client-side token bucket adds a second layer of protection. AWS SDKs track a retry budget separately from the main rate limit: once that budget depletes, the client stops retrying entirely rather than hammering a struggling service. Adaptive rate limiting takes this further by adjusting your own send rate based on live feedback: if x-ratelimit-remaining is dropping fast relative to time left in the window, slow down proactively instead of waiting for the first 429.
Pro Tip: Treat a 429 with no Retry-After header as a signal to double your current backoff multiplier, not to retry at your normal interval.
A few anti-patterns consistently make throttling worse:
- Retry storms: many workers retrying on the same fixed schedule, which recreates the original burst that triggered throttling in the first place.
- Retrying non-idempotent operations: resending a write or a payment-triggering call can duplicate side effects; only retry safely idempotent reads or calls with deduplication keys.
- Ignoring secondary limits: backing off from a primary 429 while still tripping burst limits on concurrent connections.
Our playbook on retries walks through implementation details for teams building this logic from scratch.
Polite crawling: robots.txt, crawl-delay, and site directives
Respecting a site's stated crawling rules is both a legal hedge and a practical one: compliant crawlers get blocked far less often than ones that ignore published limits.
- RFC 9309 specifies that if
robots.txtis unreachable because of a server error (5xx), crawlers must assume disallow rather than treating the absence as permission. - For other unreachable states, the same RFC allows crawlers to rely on a cached copy of
robots.txtfor up to 24 hours, which avoids re-fetching the file on every single request. - The RFC also recommends parsers impose reasonable size and complexity limits on the file itself, protecting the crawler from malformed or oversized directives.
- A
Crawl-delaydirective, while not part of the original standard, is widely honored in practice; when a site publishes one, treat it as the minimum acceptable gap between your requests to that domain.
When a site offers both a public-facing website and an authenticated API, the API is almost always the better path: it comes with documented limits instead of inferred ones, and its terms of service spell out what's permitted. Scraping the public HTML when a sanctioned API exists usually violates that same TOS, even when the technical access is trivial. Our guide to robots.txt compliance and our piece on ethical scraping practices cover this distinction in more depth, including how PII handling factors into the decision.
Proxy and header strategies: rotation trade-offs and fingerprinting
Rotating IP addresses is the most common response to persistent throttling, but the choice of proxy type carries real trade-offs.
- Datacenter proxies: cheap and fast, but their IP ranges are well cataloged, so detection systems flag them quickly, especially on sites with mature anti-bot defenses.
- Residential proxies: cost more and add latency, but route through real consumer IPs, which lowers the detection surface considerably for sites that specifically block known datacenter ranges.
- Rotation frequency: rotating too often can itself look suspicious on sites that expect session continuity; match rotation cadence to how a real user would behave on that site.
Simple IP or header rotation stops working once a target fingerprints beyond the request line. TLS fingerprinting (commonly JA3), cookie continuity, and browser feature probes can tie requests together regardless of which IP sent them, which is exactly why Cloudflare recommends session-based or fingerprint-aware rate limiting over raw IP counting on the server side. A scraper rotating IPs while reusing the same TLS stack and header order is still trivially linkable.
On the operational side, pool sizing matters more than raw proxy count: a smaller pool with healthy connection reuse often outperforms a large pool that churns connections constantly, since repeated TLS handshakes add latency and cost. Geographic targeting adds another variable, since some content varies by region and proxy location needs to match the data you're actually trying to collect. Our breakdown of CAPTCHA and session-bound detection goes further into how fingerprinting defeats naive rotation strategies.
Testing and simulation: validating retry logic safely

Burning real API quota to test whether your backoff logic works is a wasted budget. Microsoft's Dev Proxy approach, built specifically around GitHub's rate-limit behavior, simulates x-ratelimit-* headers and Retry-After responses locally, so your CI pipeline can exercise throttling logic without ever touching a production endpoint.
A solid test suite for retry handling covers:
- Missing
Retry-After: confirm your client falls back to exponential backoff instead of retrying immediately. - Secondary limit triggers: simulate a burst limit distinct from the primary limit and verify your client backs off from both independently.
- Idempotency checks: confirm retried write operations don't duplicate side effects when a response is ambiguous.
- Token-budget exhaustion: verify your client stops retrying cleanly once its retry quota depletes, rather than looping indefinitely.
Running these as local, repeatable CI tests, rather than live calls against a real API, keeps your quota intact for actual production traffic and catches regressions before they reach a target site. Our list of common scraping mistakes includes several teams have hit by skipping this step.
Monitoring, observability, and cost control
Retry logic without monitoring is a blind spot waiting to become an invoice. A handful of metrics tell you whether your scraping operation is healthy or quietly degrading:
- 429 rate: the percentage of requests hitting throttling, tracked per domain and per API key.
- Retry count per request: a rising average signals a target tightening its limits before you see outright blocks.
- Success-per-attempt ratio: how many attempts, on average, it takes to get one successful payload.
- Cost per successful extraction: the real unit economics once retries and failed attempts are priced in.
A client-side retry quota, as described in AWS's retry documentation, stops unbounded retries once the token budget depletes. That same budget concept applies directly to cost control: a hard spending cap, set before a scraping run starts, turns an unpredictable bill into a known ceiling.
Alerting should trigger on rising 429 rates before they become outright bans, not after. A dashboard tracking throttling rate alongside cost per successful payload, refreshed in near real time, catches a target tightening its limits days before your success rate collapses. Our guide to cost control pilots and our breakdown of typed failure handling both expand on building these dashboards without over-engineering them.
Gyrence approach: predictable scraping with typed failures and spending caps
We built Gyrence around the assumption that rate limits, blocks, and partial failures are normal operating conditions, not edge cases to patch around later. Every call returns a typed, discriminated-union response, including the failure cases, so an agent or pipeline can branch on exactly what happened instead of parsing error strings.
Spending caps sit at the workspace level, so a retry storm or a misconfigured crawl never turns into a surprise invoice. Our five composable primitives map directly onto safe scraping operations: Search and Map handle discovery without hammering a single endpoint, Traverse (Gyre) controls outward crawling with built-in pacing, Fetch normalizes pages to markdown, and Extract applies LLM-guided schema extraction in the same call. WebDoppler layers monitoring and webhook alerts on top, flagging drift or rising failure rates before they become a blocked account. Our retry playbook and error-handling guide go deeper on the patterns behind these primitives.
Short practical rules we follow in production
A few rules hold up across almost every scraping project we've worked through. First, Retry-After is not a suggestion: when a server sends it, honor it exactly rather than substituting your own backoff math. Second, default to backoff with full jitter, not fixed intervals, since synchronized retries recreate the exact burst that caused the throttling. Third, test retry logic against a simulator before it ever touches a live target, so the first real 429 your code sees isn't also the first one it's ever handled. Fourth, set a spending or request cap before a run starts, not after a bill arrives. None of this is exotic. It's just the difference between a scraper that degrades gracefully and one that gets an IP range permanently blocked.
**
Make scraping predictable with a managed web data API
Everything above works, but implementing retry quotas, header-aware backoff, proxy rotation, and cost dashboards from scratch is a real engineering project before you've extracted a single row of data. We built Gyrence so that work comes bundled in: typed failure responses tell you exactly why a call didn't succeed, spending caps keep a bad crawl from becoming a bad invoice, and WebDoppler watches for drift and rising throttling so you're not debugging a silent failure days later.
If your team needs deeper platform work beyond the API itself, like full pipeline builds or custom integration, Extraordinary's engineering capabilities cover that kind of end-to-end buildout. For scraping and extraction directly, our pricing page lays out the Free, Founders, Pay-As-You-Go, Standard, Growth, and Scale tiers, including Standard at $75 per month and Growth at $299 per month, so you can match a plan to your actual request volume before committing.
FAQ
What is REST API rate limiting and how is it used?
REST API rate limiting caps how many requests a client can send in a given time window, using methods like token buckets or sliding windows to count requests against an IP, API key, or user identity. Servers signal the limit through headers like x-ratelimit-remaining and return an HTTP 429 status code once the cap is exceeded, often with a Retry-After header telling the client when to try again.
Is AI scraping illegal?
Scraping itself isn't inherently illegal, but legality depends on what's scraped, how it's accessed, and what a site's terms of service permit. Respecting robots.txt directives, avoiding authentication bypass, and handling personal data carefully are the practical safeguards; our ethical scraping guidance covers the main considerations in more detail.
Does Twitter allow data scraping?
Social platforms generally restrict scraping through their terms of service and enforce limits through authenticated API access instead. The safest path on any platform with a sanctioned API is to use that API under its published rate limits rather than scraping the public site directly.
Is HTTP 429 a rate limit?
Yes, HTTP 429 specifically means "Too Many Requests", indicating the client has exceeded a rate limit for that endpoint or resource. RFC 6585 formalized this status code and notes that 429 responses must not be cached, since the condition is tied to that specific client's request history.
Sources
- Retry behavior
- 429 Too Many Requests - HTTP | MDN
- RFC 6585 - Additional HTTP Status Codes
- What is rate limiting? (Cloudflare)
- How to test GitHub API rate limit handling

