Blog
Discover our latest articles and blogs

Polite Crawling: Developer Defaults for 429s, Crawl Caps, and Audits
Build a polite crawler with bot identification, robots.txt rules, per host pacing, and 429 or 5xx backoff. Set crawl limits and track failures for audits.

From WHATWG to Go: Content Type Detection for Developers
A developer focused operational guide that maps WHATWG rules to platform behaviors. Learn the 512 byte signature limit, signature versus parse checks, and...

Cut Blocking Risk: 20s Backoff + Spending Caps on Scraping Rate Limits
Developer playbook for scraping rate limits: read 429/Retry-After, use backoff with jitter, test retries in CI, and enforce spending caps to avoid blocks.

When Sitemaps Hit 50,000 URLs: Engineer Fixes to Restore Crawling
Practical checklist for engineers: when sitemaps matter, the 50,000 URL/50 MB limits to watch, and the crawl-pipeline fixes that actually restore indexing.

Prevent Lost Pages: 5 Validation Steps for Canonical URLs (Engineers)
Validate canonical URLs: a developer playbook. Run Link header checks first, confirm targets return 200, resolve URLs, and log provenance.

Engineers: Preserve HTML in Content Normalization for RAG Scraping
Roadmap for engineers to normalize content: preserve HTML structure, keep structured and plain text outputs, validate on RAG/LLM tasks.

Five minute developer checklist: Headless vs HTTP scraping, HTTP first
A developer-first checklist to choose HTTP-first scraping or headless browsers. Run a five-minute DevTools test, compare costs and failure modes, and see...

Fix Scraping API Errors for Developers: 3 Checks and Typed Failures
Engineering-first triage for scraping API errors: run 3 quick checks, apply safe retries with backoff, and use typed failures and billing caps.

Engineers Save Crawl Bandwidth with URL Deduplication and Gyrence
Implementation-first guide for engineers on URL deduplication: canonicalize requests, use request fingerprints, Bloom filters, simhash, ops metrics,...