Blog

Discover our latest articles and blogs

Polite Crawling: Developer Defaults for 429s, Crawl Caps, and Audits
October 10, 2026

Polite Crawling: Developer Defaults for 429s, Crawl Caps, and Audits

Build a polite crawler with bot identification, robots.txt rules, per host pacing, and 429 or 5xx backoff. Set crawl limits and track failures for audits.

From WHATWG to Go: Content Type Detection for Developers
October 9, 2026

From WHATWG to Go: Content Type Detection for Developers

A developer focused operational guide that maps WHATWG rules to platform behaviors. Learn the 512 byte signature limit, signature versus parse checks, and...

Cut Blocking Risk: 20s Backoff + Spending Caps on Scraping Rate Limits
October 8, 2026

Cut Blocking Risk: 20s Backoff + Spending Caps on Scraping Rate Limits

Developer playbook for scraping rate limits: read 429/Retry-After, use backoff with jitter, test retries in CI, and enforce spending caps to avoid blocks.

When Sitemaps Hit 50,000 URLs: Engineer Fixes to Restore Crawling
October 7, 2026

When Sitemaps Hit 50,000 URLs: Engineer Fixes to Restore Crawling

Practical checklist for engineers: when sitemaps matter, the 50,000 URL/50 MB limits to watch, and the crawl-pipeline fixes that actually restore indexing.

Prevent Lost Pages: 5 Validation Steps for Canonical URLs (Engineers)
October 6, 2026

Prevent Lost Pages: 5 Validation Steps for Canonical URLs (Engineers)

Validate canonical URLs: a developer playbook. Run Link header checks first, confirm targets return 200, resolve URLs, and log provenance.

Engineers: Preserve HTML in Content Normalization for RAG Scraping
October 5, 2026

Engineers: Preserve HTML in Content Normalization for RAG Scraping

Roadmap for engineers to normalize content: preserve HTML structure, keep structured and plain text outputs, validate on RAG/LLM tasks.

Five minute developer checklist: Headless vs HTTP scraping, HTTP first
October 4, 2026

Five minute developer checklist: Headless vs HTTP scraping, HTTP first

A developer-first checklist to choose HTTP-first scraping or headless browsers. Run a five-minute DevTools test, compare costs and failure modes, and see...

Fix Scraping API Errors for Developers: 3 Checks and Typed Failures
October 3, 2026

Fix Scraping API Errors for Developers: 3 Checks and Typed Failures

Engineering-first triage for scraping API errors: run 3 quick checks, apply safe retries with backoff, and use typed failures and billing caps.

Engineers Save Crawl Bandwidth with URL Deduplication and Gyrence
October 2, 2026

Engineers Save Crawl Bandwidth with URL Deduplication and Gyrence

Implementation-first guide for engineers on URL deduplication: canonicalize requests, use request fingerprints, Bloom filters, simhash, ops metrics,...