← Back to blog

From WHATWG to Go: Content Type Detection for Developers

October 9, 2026
From WHATWG to Go: Content Type Detection for Developers

For production use, treat Content-Type and filename as advisory. Rely on a short, byte-level signature check plus targeted parsing for text formats before trusting what a file claims to be. When the signature is unrecognized and parsing fails, fall back to application/octet-stream and store the file safely rather than guessing.


TL;DR:

  • Use headers and extensions only for routing or display; enforce an allowlist with signatures for binaries and full parsing for JSON, CSV, and XML.
  • Limit signature checks to 512 bytes; because formats such as DOCX and XLSX use ZIP containers, inspect archive structure read only without fully extracting files.
  • Store uploads outside the web root, decode and encode images again, and serve them with a nosniff header; scan or sandbox high risk files.
  • Python’s mimetypes module checks extensions only, while Go’s detector reads up to 512 bytes; neither should serve as a standalone security gate.

Gyrence
Build More Reliable Web Data Pipelines
Gyrence provides typed web data APIs for search, fetching, extraction, and site traversal, with explicit failure modes for agent workflows.
Explore Gyrence

Table of Contents

Choosing the right content type detection method

Different situations call for different levels of rigor. Picking the wrong one is how security bugs and broken renders end up in production.

  • Content-Type header: fast and useful for routing or UX decisions, but trivial to spoof, so never use it for security.
  • File extension mapping: the cheapest signal, acceptable for best-effort features like icon selection, worthless as a trust boundary.
  • Magic bytes and signature checks: reliable for most binary formats and the right default when accepting user uploads.
  • Sniffing algorithms: good for browser compatibility, risky when a wrong classification could let a file execute as script.
  • Full parse and validation: the only dependable approach for text formats (JSON, XML, CSV) that carry no byte signature at all.

A production pipeline usually layers these: header for routing, signature for binaries, parse for text, with the stricter check winning whenever they disagree.

How magic bytes and signature detection actually work

Magic bytes are fixed byte sequences near the start of a file that identify its format. A PNG begins with 89 50 4E 47, a JPEG with FF D8 FF, a GIF with GIF87a or GIF89a, a ZIP archive with 50 4B 03 04, and a WebAssembly binary with 00 61 73 6D. Checking these bytes is fast, deterministic, and doesn't require parsing the whole file.

The hard part is container ambiguity. ZIP is the backing format for .docx, .xlsx, .jar, and plenty of other types, so a matching ZIP signature only tells you the outer shell. You need to peek inside the archive's internal structure (such as looking for [Content_Types].xml in Office formats) without fully extracting an untrusted archive.

  • Text formats like JSON, CSV, and plain HTML have no magic bytes at all, so signature checks can't help there.
  • Polyglot files, crafted to be valid in two formats simultaneously, can slip past a naive signature check and a naive parser alike.
  • Reading more than a small header window wastes time and widens your exposure to oversized or malicious input.

Pro Tip: Cap your signature read at 512 bytes and treat any inner-archive probing as read-only, never executing or fully decompressing content you haven't sandboxed.

What the sniffing standards actually say

Browsers don't just trust the Content-Type header, and the rules for when they override it are formally specified. The WHATWG MIME Sniffing Standard defines the algorithm user agents use, built around reading at most the first 512 bytes of a resource and matching them against a table of known patterns. That 512-byte ceiling exists specifically to bound latency and avoid letting an attacker force excessive reads.

Browsers can override a server's declared type under specific conditions, which is exactly why X-Content-Type-Options: nosniff exists: it tells the browser to respect the declared header and skip sniffing entirely, as documented in MDN's media types guide.

  • Sniffing exists because many servers historically misreport Content-Type, so clients compensate.
  • Without nosniff, a user-uploaded file misclassified as HTML or SVG can execute as script in the browser, an XSS path that has shown up in real-world upload features.
  • Different agents implement different heuristics, so sniffing behavior isn't perfectly portable across browsers or tools.

Setting nosniff on anything you serve back to users, especially uploaded content, closes off a meaningful attack surface for close to zero engineering cost.

Platform libraries: what each one actually checks

Knowing what a given API inspects, bytes or just a string, determines whether you can trust its output for a security decision.

  • Go's net/http.DetectContentType follows the WHATWG sniffing algorithm, reads up to 512 bytes, and falls back to application/octet-stream when nothing matches, according to the Go standard library documentation. It's suitable for compatibility checks, not a standalone security gate.
  • Python's mimetypes module maps extensions to types using static tables and performs no byte inspection at all, per the Python documentation. Treat it as a convenience lookup, never a validator.
  • The Unix file command runs a magic database plus several heuristics, including regex matching for text formats, as described in its man page. It's a solid tool for CI pipelines and offline inspection, less practical for inline request handling.
  • Third-party libraries across Rust, Node, and .NET vary widely in signature coverage and maintenance activity: check how recently the signature table was updated before depending on one in production.

Security checklist for handling user-supplied files

The OWASP File Upload Cheat Sheet is the reference point here, and its core guidance translates into a short operational checklist.

  1. Never use the client-supplied Content-Type header or filename extension as the basis for a security decision.
  2. Validate against an allowlist of accepted types rather than trying to blocklist dangerous ones.
  3. Run a signature check on binary uploads and a full parse or validation pass on text formats before accepting them.
  4. Store uploaded files outside the web root, and re-encode images (decode then re-encode) to strip embedded payloads and polyglot content, a mitigation OWASP specifically recommends.
  5. Set X-Content-Type-Options: nosniff and an appropriate Content-Disposition header on anything you serve back.
  6. Route high-risk uploads through malware scanning or a sandboxed environment before they touch the rest of your system.

A safe detection pipeline you can implement today

A pipeline that holds up in production looks like this: treat the advisory header as a hint, run a bounded signature check (512 bytes or fewer), parse or validate text formats directly, normalize and re-encode where applicable, then store and serve through the checklist above.

  • Tag each detection result with a confidence level (header-only, signature-confirmed, parse-validated) so downstream code knows how much to trust it.
  • Surface failure modes explicitly, such as "signature unmatched" or "parse failed," rather than silently defaulting to a guess.
  • For ingestion pipelines feeding retrieval-augmented generation or other LLM workflows, a confidence tag lets downstream systems decide whether re-validation is worth the cost before processing untrusted content.

Pro Tip: Log the detected type alongside the declared type for every upload. Divergence between them is one of the cheapest signals you can collect for catching spoofed files early.

Where detection pipelines usually go wrong

Where detection pipelines usually go wrong — overview diagram

The trade-off between compatibility and security is real: strict signature checking will reject some legitimate files that sniffing would have accepted, and that's usually the right default to choose.

The recurring mistakes are predictable: trusting an extension because it "looked fine," skipping sandboxing because the upload volume was low, or forgetting nosniff on a response path added after the original security review. Track misclassification rate and file-related incident count as your two honest success metrics.

— Glen

Content-type-aware ingestion without the guesswork

Detection logic like the pipeline above is exactly the kind of plumbing that's easy to get half right and expensive to get wrong at scale. Our API handles content-type-aware routing natively: every call to Gyrence returns a typed, discriminated-union response that tells you what we found and how, including the failure cases, instead of forcing you to guess when a fetch comes back ambiguous.

Gyrence

For teams that need to watch pages for content or type drift over time rather than fetching once, WebDoppler adds monitoring with webhook alerts on top of the same detection logic. Check our pricing to see which plan fits your call volume, with the Standard tier starting at the published monthly price.

FAQ

Can you give me an example of a MIME type?

A MIME type follows a type/subtype structure, such as image/png, application/json, or text/html, sometimes with parameters like text/html; charset=utf-8. The structure and registration rules for these strings are defined in RFC 6838.

How do I open a MIME file?

There's no single "MIME file" format: MIME types describe the content type of a file or data stream, not a container you open directly. Open the underlying file with whatever application matches its actual type, such as a browser for text/html or an image viewer for image/jpeg.

How do I identify a MIME type?

Read a small header of the file (512 bytes or fewer is standard) and compare it against known magic byte signatures, which is how tools like Go's net/http.DetectContentType and the Unix file command work. For formats without signatures, such as JSON or CSV, parse or validate the content directly instead.

How do I fix a MIME type mismatch?

First check whether the server is sending an incorrect Content-Type header for the file it's serving, which is a configuration fix on the server side. If the file's actual bytes don't match its claimed type at all, re-run signature detection and correct the stored type, and consider setting X-Content-Type-Options: nosniff so browsers stop guessing around the mismatch.

Sources