Start by canonicalizing requests and using a request fingerprint. This single step catches most duplicate crawl traffic before it wastes a fetch. Add content fingerprints like simhash or probabilistic structures like Bloom filters only when scale or heavy content variation demands them, since each adds memory cost and a small, tunable false-positive rate in exchange for real bandwidth savings.
TL;DR:
- URL normalization and canonicalization are essential to prevent silent duplicate crawling caused by trivial differences like query parameter order or default ports.
- Building a persistent deduplication index across multiple crawls reduces unnecessary fetches, saving crawl budget and server resources.
- Content fingerprinting with simhash effectively detects near-duplicate pages, such as variant article templates, avoiding redundant storage and indexing.
- Bloom filters provide scalable, memory-efficient URL deduplication with a tunable false-positive rate, suitable for large-scale crawling environments.
- Implementing layered deduplication—canonicalize URLs, request fingerprints, Bloom filters, and content similarity—maximizes efficiency and reliability in large or ongoing crawls.
Table of Contents
- 1. URL-based deduplication within a single crawl
- 2. Deduplication across multiple crawls and dedupe indexes
- 3. Content-based and near-duplicate detection with simhash
- 4. Probabilistic structures for scalable URL deduplication
- 5. URL normalization and canonicalization best practices
- 6. Practical implementation patterns and an operations checklist
- How Gyrence applies these techniques
- Picking defaults for correctness versus cost
- Gyrence as the operationally simpler path
- Sources
- FAQ
1. URL-based deduplication within a single crawl
Within a single crawl run, the fastest win is a request fingerprint built from a canonicalized URL, the HTTP method, and the request body. Canonicalization matters because example.com/page?b=2&a=1 and example.com/page?a=1&b=2 are the same resource, but a naive string comparison treats them as different. Fragments should be dropped entirely since they never reach the server, and query parameter order should be normalized before hashing.

Scrapy's RFPDupeFilter is a working reference: it computes a SHA-1 hash over the canonicalized URL, method, and body, using w3lib.url.canonicalize_url under the hood, as documented in the Scrapy request and response docs. Fragments are ignored and query order is normalized by default, which is why it holds up across large crawls without extra configuration.
A robust fingerprint pipeline generally includes:
- Lowercase the scheme and host, strip default ports.
- Percent-decode safe characters and re-encode consistently.
- Sort query parameters unless order is functionally significant.
- Hash the normalized URL plus method and body with SHA-1 or a similar fixed-length digest.
- Persist seen fingerprints to disk (a jobdir file, for instance) when a crawl may resume after interruption.
2. Deduplication across multiple crawls and dedupe indexes
A single run's dedupe filter forgets everything once the process exits. If you re-crawl a site weekly, you need a persistent dedupe index that survives across jobs, otherwise every run refetches URLs you already have. This matters most for large sites where re-fetching unchanged pages burns both your crawl budget and the target server's goodwill.
Building that index well means handling a few recurring problems:
- Choose a persistent store, whether a key-value database, a compact on-disk set, or a managed service, sized for your expected URL volume.
- Commit fingerprints in batches so a crash mid-run does not leave the index half-written or duplicated.
- Support merge operations when parallel workers write to the same index concurrently.
- Expire or compact entries on a schedule so the index does not grow unbounded as content churns.
- Decide between a centralized dedupe service, useful when many jobs share state, and per-job jobdir files, which are simpler but isolate each run.
Interrupted runs are the real test of this design. A dedupe index that cannot roll back a partial commit will eventually mark a genuinely new URL as seen, and you lose that page silently.
3. Content-based and near-duplicate detection with simhash
URL-level dedupe catches identical requests, but it misses pages that are different URLs with near-identical content: printer-friendly versions, session-tagged mirrors, or syndicated articles with a different header. That's where content fingerprinting earns its cost. Charikar's simhash algorithm, described in the WWW 2007 paper on near-duplicate detection, reduces a page's feature vector into a small, fixed-width fingerprint where similar pages produce fingerprints that differ in only a few bits.
Building one in practice looks like this:
- Extract shingles or weighted tokens from the page's meaningful text, excluding boilerplate like navigation and footers.
- Hash each feature and combine the hashes into a single f-bit fingerprint, typically 64 bits.
- Compare fingerprints using Hamming distance rather than exact match.
- Use a small threshold, around k=3, which the source paper notes works well for repositories scaling into billions of pages.
- Run detection in batch for large backfills, and online per-page only when the crawl rate can absorb the extra computation.
Near-duplicate detection pays off clearly when a site serves the same article through multiple template variants. Catching that at the content level avoids storing and re-indexing the same information under five different URLs.
4. Probabilistic structures for scalable URL deduplication
Once a crawl's seen-set grows past what fits comfortably in memory as an exact structure, Bloom filters become the standard tool. A 2016 conference paper on Bloom filter application in web crawlers frames this directly: Bloom filters trade a small, tunable false-positive rate for a large reduction in memory footprint compared to storing full URLs or hashes.
Sizing one means picking bit-array size m and hash count k for an expected number of items n, since the false-positive rate r is a direct function of those three values. Non-cryptographic hash functions like Murmur or Jenkins are common here because speed matters more than collision resistance for this use case.
- Size
mandkagainst your expectedn, not your current crawl size, since undersizing degrades the false-positive rate as the filter fills. - Prefer Murmur or Jenkins hashes over cryptographic hashes for speed at scale.
- Sample a fraction of filter-suppressed URLs and re-check them against an exact source, such as a content hash, to measure real-world false-positive rate.
- Consider a counting Bloom filter when you need to support deletions or expiration.
Pro Tip: Run your false-positive sampling check continuously in production, not just during initial tuning. Filter behavior drifts as it fills.
Bloom filters are the right choice when memory is the binding constraint and a small error rate is tolerable. When correctness must be exact, such as billing-relevant dedupe, an exact store is worth the extra memory.
5. URL normalization and canonicalization best practices
Normalization done before fingerprinting eliminates a large share of duplicates without any hashing at all. The steps are mechanical but easy to get wrong in a way that silently inflates your crawl:
- Lowercase the scheme and host; strip default ports (
:80for HTTP,:443for HTTPS). - Percent-decode unreserved characters and normalize path segments, collapsing
./and../. - Sort or strip query parameters, dropping known tracking parameters like
utm_sourcethat don't change page content. - Drop fragments unless your crawler specifically needs client-side routing state.
- Respect server-declared canonical signals, whether a
rel=canonicaltag or an HTTPLinkheader, as specified in RFC 6596.
RFC 6596 also warns about chained canonicals, where page A points to B, which points to C. A crawler that follows blindly can end up discarding unique content if the chain terminates at a 404 or an unrelated page. Verify that a canonical target actually returns content before trusting it.
6. Practical implementation patterns and an operations checklist
A layered pipeline works better than any single technique alone: canonicalize the URL, check a request fingerprint for exact within-run dupes, check a Bloom filter for cross-run seen state, and only fall back to simhash comparison for pages that pass both and warrant content-level scrutiny.
url = canonicalize(raw_url)
fp = sha1(url + method + body)
if fp in seen_requests: skip()
if bloom.might_contain(url): sample_check(url)
else: bloom.add(url); fetch(url)
if content_dedupe_enabled:
sh = simhash(extract_features(page))
if hamming(sh, index.nearest(sh)) <= k: skip_as_near_duplicate()
Export these metrics from any production dedupe layer:
- Dupe rate: the percentage of requests filtered before fetch.
- False-positive sample rate from your Bloom filter sampling checks.
- Memory pressure on the seen-set or filter structure.
- Dedupe-index growth rate over time, to catch runaway URL spaces early.
Statistic Callout: Bloom filter architectures reduce memory use versus storing full URL lists, at the cost of a tunable false-positive rate that teams typically measure through sampling. That trade-off is the whole design decision in one line.
Before enabling any of this in production, test settings against a sampled subset of URLs and compare the dupe rate against a manual spot-check. A checklist worth following: canonicalize first, fingerprint second, size your Bloom filter for expected scale not current scale, sample continuously, and never disable dedupe globally when you only mean to allow one request through, since Scrapy's own DUPEFILTER_CLASS guidance recommends dont_filter=True on individual requests instead.
How Gyrence applies these techniques
Gyrence's Traverse primitive maps a site's URL graph and applies normalization before a page ever gets queued, which cuts the fingerprinting workload downstream. Fetch and Extract return typed, discriminated-union responses, including failure cases, so a dedupe-heavy job surfaces a bad canonical target or a fetch error instead of hiding it. Spending caps mean a runaway crawl frontier, the kind that inflates a Bloom filter faster than planned, never turns into a surprise bill.

Picking defaults for correctness versus cost
Exactness matters when a duplicate costs you money or a broken report. Otherwise, a small false-positive rate is a fair trade for memory savings. Small teams should start with request fingerprints and a persisted jobdir. Start there today.
— Glen
Gyrence as the operationally simpler path
Building and tuning fingerprint filters, Bloom filters, and simhash pipelines is real engineering time that most teams would rather spend on their product. Gyrence's five composable primitives, Search, Traverse, Fetch, Extract, and Map, handle URL mapping and page fetching with normalization already applied, and WebDoppler adds monitoring and webhook alerts on top for tracking content drift over time.
Pricing runs on predictable, credit-based plans listed on the Gyrence pricing page, so a dedupe-heavy crawl job doesn't turn into an unpredictable invoice. Check WebDoppler if ongoing monitoring is part of your workflow, or review plans on the pricing page to see what fits your crawl volume.
Sources
RFC 6596 covers the canonical link relation. The WWW 2007 simhash paper covers near-duplicate detection algorithms. Scrapy's docs show a concrete fingerprint implementation. The Bloom filter paper covers sizing and false-positive trade-offs at crawl scale.
For related engineering context, see how bulk domain crawling handles frontier design at scale, how HTTP caching and a polite fetcher reduce redundant network load, and a case study on stopping an agent from looping on paginated content.
- Scrapy documentation — request/response topics
- Detecting Near-Duplicates for Web Crawling (WWW 2007)
- Application of Bloom Filter for Duplicate URL Detection in a Web Crawler (2016 conference paper abstract)
- RFC 6596 — The Canonical Link Relation
FAQ
How can I check if my website is crawlable?
Check your robots.txt file for disallowed paths and confirm your sitemap lists the URLs you expect indexed. Fetching a page with a bare HTTP client and comparing the response to what a browser renders will reveal if JavaScript rendering is blocking crawler access.
Can you give me an example of a web crawler?
Scrapy is a widely used open-source crawling framework with built-in request deduplication through its RFPDupeFilter. Managed services like Gyrence offer crawling as an API primitive, called Traverse, for teams that want URL mapping without operating the crawler infrastructure themselves.
What does it mean when a website is crawled?
Crawling means an automated program fetches a page's content, follows its links, and repeats the process across a site or the web. The results typically feed into search indexing, data extraction, or monitoring systems.
How can I force Google to crawl my website?
You can submit a URL or sitemap directly through Google Search Console to request indexing. Ensuring your robots.txt and canonical tags aren't accidentally blocking pages, per RFC 6596 guidance, also helps crawlers reach the right content.
Why does URL deduplication matter for crawl efficiency?
Deduplication avoids wasting fetch requests, storage, and processing time on content you've already collected. Layered approaches, combining request fingerprints, Bloom filters, and near-duplicate detection, keep large crawls efficient without exact-match memory costs at every step.

