← Back to blog

Polite Crawling: Developer Defaults for 429s, Crawl Caps, and Audits

October 10, 2026
Polite Crawling: Developer Defaults for 429s, Crawl Caps, and Audits

Polite crawling means fetching web pages in a way that identifies your bot, obeys robots.txt, paces requests to avoid overloading servers, and backs off the moment a site signals distress. The immediate rules: publish a real User-Agent string, check robots rules before every fetch, throttle per host, and treat 429 or 5xx responses as a command to slow down, not retry harder.


TL;DR:

  • Publish a descriptive User Agent with documentation and contact details, honor cache directives and 304 responses, and never bypass authentication walls or CAPTCHAs.
  • Cache robots.txt for 24 hours, follow no more than five redirects, and treat server errors as undefined permissions rather than open access.
  • Start at 1 second for large sites, 5 for medium sites, and 10 to 15 for small hosts; reduce requests if p95 latency doubles.
  • For 429 or 5xx responses, cut concurrency first, honor the server's stated wait, then add jittered exponential backoff; do not retry other 4xx errors.
  • Alert when 5xx responses exceed 1% over 10 minutes or host TTFB doubles, then throttle that host and pause crawling if the pattern persists.

Gyrence
Build with Structured Web Data
For crawl workflows that need structured web data, Gyrence offers composable API primitives for searching, fetching, extracting, and mapping sites.
Explore Gyrence

Table of Contents

Core politeness rules and why each matters

A polite crawler behaves like a predictable guest, not an unannounced flood of traffic. Each rule below exists to prevent a specific failure: blocked IPs, angry site owners, or a server that falls over under load.

  • Send a descriptive User-Agent string with a link to a public crawler doc and a contact method, so site owners can identify and reach you instead of guessing.
  • Check robots.txt and X-Robots-Tag before fetching, and cache the result for up to 24 hours rather than re-fetching on every request.
  • Respect cache directives and 304 responses: if a page hasn't changed, don't re-download it.
  • Use per-host queues with limited parallel connections, and never attempt to bypass authentication walls or CAPTCHAs.

These aren't courtesy gestures. A site that sees a clean User-Agent and sane pacing is far less likely to block your range outright, which keeps your crawl running instead of failing silently. For a deeper walkthrough of rate limits and PII handling alongside these basics, see our guide on ethical web scraping practices.

Robots Exclusion Protocol: parsing, redirects, and caching fallbacks

RFC 9309 formalizes what used to be a loose convention into an actual protocol, and the details matter more than most implementations admit.

  • Match the most specific user-agent group that applies to your crawler; when multiple rules could apply, the longest matching path rule wins, not the first one listed.
  • Follow up to five redirects when fetching /robots.txt; beyond that, treat the file as unreachable.
  • Cache robots.txt for about a day, and if the file is unreachable due to a server error, treat permissions as undefined rather than assuming full access.
  • Apply a parse-size limit of at least several hundred kilobytes and ignore content beyond it rather than attempting to parse malformed or oversized files.

Robots.txt is a public file, not a security boundary: listing sensitive paths there can expose them to anyone reading the file, so access control still belongs on the server side. Our breakdown of how a single server error can block your entire crawl covers the fallback behavior in more detail.

HTTP signals and backoff strategies for 429 and 5xx responses

Status codes aren't just error reporting, they're instructions. RFC 6585 defines 429 Too Many Requests specifically so servers can tell clients to slow down, often with a Retry-After header stating exactly how long to wait.

  • Treat 429 and any 5xx response as a signal to reduce request rate immediately, not a prompt to retry on the same schedule.
  • Treat other 4xx codes (404, 403) as "don't retry": the problem is the request, not the pace.
  • Parse and honor Retry-After when present; when absent, fall back to your own backoff schedule.
  • Apply backoff in order: cut concurrency first, then add exponential backoff with jitter, then cap the maximum wait so retries don't stall indefinitely.

Pro Tip: Reduce concurrency before you touch your delay timer. Dropping active connections frees server resources immediately, while a longer sleep interval only helps on the next request.

Our retry and backoff playbook walks through working examples of this ordering in practice.

How to pick per-host delays and adapt pacing over time

There's no universal crawl delay, but there are reasonable starting points depending on what you're crawling.

  1. Start with 1 second between requests for large, well-provisioned sites, 5 seconds for mid-sized sites, and 10 to 15 seconds for small or self-hosted sites where you can't confirm infrastructure.
  2. Measure median response time and p95 latency on an ongoing basis as you crawl.
  3. If p95 latency doubles compared to your baseline, cut your request rate meaningfully rather than waiting for errors to appear.
  4. Implement a token-bucket or leaky-bucket limiter per host, not a single global limiter, so one slow site doesn't starve or get flooded disproportionately.
  5. Prefer per-host queues over global parallelism: ten threads spread across ten hosts is safer than ten threads hammering one.

Google's crawl budget documentation confirms this dynamic: crawl capacity adjusts based on server responsiveness, so faster, healthier responses earn more headroom over time, and persistent latency shrinks it.

Metrics, alerts, and warning signs your crawler is causing harm

A polite crawler watches its own impact, not just its own throughput. Track error rate, median and p95 latency, connection failures, and spikes in 429 responses as your core health signals.

Google's own crawling guidance confirms that persistent 5xx and 429 responses cause search crawlers to slow down automatically, and that 304 responses save origin resources by letting crawlers skip re-fetching unchanged content. The same logic should drive your own thresholds.

  • Alert when sustained 5xx responses exceed 1% of requests over a 10-minute window.
  • Alert when TTFB doubles relative to your rolling baseline for that host.
  • Alert on any sharp rise in 429 responses, even if the overall error rate still looks low.

When thresholds trip, the response should be automatic: throttle the offending host, pause the crawl entirely if the pattern persists, and notify the contact listed in your own crawler documentation so a human can review it.

Preflight and runtime checklist for a compliance-friendly crawl

Before any crawl touches production traffic, work through a short list. This isn't bureaucracy, it's what keeps a crawl from becoming an incident.

  1. Publish a crawler documentation page: User-Agent string, contact email, and client IP ranges if you operate from a known range.
  2. Implement robots.txt fetch and parse logic with 24-hour caching and a defined parse-size limit.
  3. Configure per-host queues, a token-bucket rate limiter, and a retry/backoff policy tied to status codes.
  4. Set a maximum parallelism ceiling per host and globally.
  5. Enable logging and audit trails for every request, response code, and retry decision.
  6. Set spending caps on crawl budget or API credits before any large run.
  7. Run a small-scale test crawl against a representative host to calibrate delays before scaling up.

Pro Tip: Treat your test crawl as a calibration step, not a formality. A 50-page test run against a real host tells you more about safe pacing than any default value borrowed from documentation.

Our AI agent web browsing checklist and crawl-depth limits guide both expand on these defaults for agent-driven crawling specifically.

How Gyrence approaches safe crawling defaults

We built our primitives around the assumption that a crawler should fail loudly and predictably, not silently overload a host or blow through a budget. Search, Traverse, Fetch, Extract, and Map each return a typed, discriminated-union response, including failure cases, so your code can react to a rate limit or a blocked path instead of guessing why a request came back empty.

  • Crawl-depth limits on Traverse prevent a single starting URL from spiraling into an uncontrolled site-wide fetch.
  • Spending caps at the workspace level stop a misconfigured crawl from generating a surprise bill or DDoS-like request volume.
  • Structured failure modes surface robots blocks, rate limits, and server errors as distinct, typed outcomes rather than generic exceptions.

For teams building this logic from scratch, our crawl-depth limits guide and AI agent web browsing checklist cover the same defaults in more implementation detail.

What actually matters in polite crawling

Most polite-crawling advice treats robots.txt compliance as the finish line. It's the floor, not the ceiling. A crawler can obey every disallow rule perfectly and still take a small site offline by hammering it with twenty concurrent connections on a path robots.txt never mentioned.

Robots compliance does not prevent overload

The conventional checklist approach, set a crawl delay, check robots, call it done, misses the part that actually prevents harm: watching server behavior in real time and reacting to it. A fixed delay chosen once and never revisited is a guess dressed up as a policy. The sites that get hurt by crawlers aren't usually the ones with strict robots.txt files; they're the ones running on modest infrastructure that a well-meaning crawler never bothered to measure.

If you prioritize one thing from everything above, prioritize adaptive pacing over static rules. A token-bucket limiter that tightens when p95 latency climbs will protect a host that a flat "5 seconds between requests" policy would still overwhelm during a traffic spike. Static rules are easy to write and easy to audit, which is exactly why they're overused. Reacting to the server's actual state is harder to implement and far more effective.

— Glen

Where Gyrence fits if you'd rather not build this yourself

Everything above describes what you'd need to implement by hand: User-Agent documentation, robots parsing with caching and fallback logic, per-host rate limiters, backoff policies, and spending controls to keep a crawl from running away from you. Our API bundles those controls into five composable primitives, Search, Traverse, Fetch, Extract, and Map, plus a hosted MCP endpoint, so the politeness logic is already built into every call rather than something your team maintains separately; for deeper insight into optimizing discoverability and technical SEO, see this case study on SEO improvement after a national relaunch.

Gyrence

Spending caps can help prevent a runaway crawl from generating a surprise invoice, and responses may include failure cases to help your code understand why a fetch did not succeed. For ongoing visibility into crawl behavior and site changes over time, WebDoppler monitors and alerts on drift without requiring you to build that layer yourself.

If you want to see how this holds up against a real site, our pricing page lists the Free and Pay-As-You-Go tiers, both suited to running a calibration crawl before you commit to anything larger.

Where Gyrence fits if you'd rather not build this yourself — overview diagram

FAQ

What is polite crawling?

Polite crawling is the practice of fetching web pages in a way that identifies your bot, obeys robots.txt rules, paces requests to avoid overloading servers, and backs off when a site returns error or rate-limit signals. It prioritizes the target server's stability over raw crawl speed.

What is crawling?

Crawling is the automated process of fetching web pages, usually by following links outward from a starting URL, to discover and collect content at scale. Search engines and data pipelines use crawling to build an index or dataset without manually visiting each page.

What is crawling vs scraping?

Crawling refers to discovering and fetching pages, typically by following links across a site or the web. Scraping refers to extracting specific structured data from the pages once fetched, so a single pipeline often crawls first and scrapes second.

How does Google crawler see my site?

Google's crawler evaluates server responsiveness continuously and adjusts its crawl rate based on it: persistent 5xx or 429 responses cause it to slow down, while fast, healthy responses allow more requests over time. Returning 304 (Not Modified) for unchanged pages also reduces load and helps preserve crawl capacity.

How do I choose a safe crawl delay for a new site?

A reasonable starting point is 1 second for large, well-provisioned sites and 10 to 15 seconds for small or self-hosted ones, adjusted based on measured response times rather than left static. Google's crawl budget guidance confirms that crawl capacity itself is dynamic and tied to how the server responds over time.

Sources