← Back to blog

Pydantic Web Extraction With Gyrence, No LLM Charge for Developers

September 27, 2026
Pydantic Web Extraction With Gyrence, No LLM Charge for Developers

Pydantic web extraction means declaring a typed Pydantic model, emitting its JSON Schema, and using that schema to drive a structured or LLM-assisted extractor that returns validated JSON. You get model_validate round-tripping, field-level errors instead of silent garbage, and a contract downstream systems can trust. It's the right call for messy, inconsistent markup. It's the wrong call for stable, high-volume pages where a CSS selector already does the job for less money.


TL;DR:

  • Pydantic schema extraction is ideal for inconsistent markup, multi-site coverage, and early-stage schema discovery, despite higher costs and latency.
  • Selector-based extraction is more cost-effective and faster for stable sites with predictable layouts used at high volume.
  • Building clear, descriptive fields and using nested models improves extraction accuracy and reduces validation noise across diverse web pages.
  • Effective preprocessing, like removing noise and using scope selectors, enhances extraction reliability and reduces errors.
  • Handling partial or inconsistent data requires separating fetch failures from validation errors, and normalizing data post-extraction to maintain schema stability.

Gyrence
Turn Web Pages Into Typed Data
Gyrence’s Extract primitive turns web pages into structured JSON using prompts or schemas, with typed responses that include failure cases.
Explore Gyrence

Table of Contents

Should You Use Pydantic Schema Extraction or CSS Selectors?

The decision comes down to how much the markup moves and how many sites you're covering.

Schema-guided extraction wins when you don't control the source, layouts vary across domains, or you're prototyping and don't yet know what a stable selector would even look like. You define the shape you want. The extractor figures out where on the page that shape lives. Selectors win when the DOM is predictable and you're running the same extraction thousands of times a day, because a selector costs almost nothing per call while an LLM-guided extraction carries real latency and token cost.

  • Prefer Pydantic-guided extraction for unknown or inconsistent markup, multi-site coverage, and early-stage schema discovery.
  • Prefer selectors for a single stable site, high call volume, and when the layout hasn't changed in months.
  • Consider hybrid: discover with a model, freeze into cached selectors once the layout proves stable.

A Minimal Pydantic Web Extraction Pattern

Here's the skeleton most production pipelines converge on: declare the model, emit its schema, call an extractor, validate what comes back.

Four-step Pydantic extraction validation flow

from pydantic import BaseModel, Field, ValidationError
from typing import Optional

class Product(BaseModel):
    name: str = Field(..., description="Product title as shown on the page")
    price_usd: Optional[float] = Field(None, description="Current sale price in USD, not list price")
    in_stock: Optional[bool] = Field(None, description="True if the page shows an add-to-cart option")

schema = Product.model_json_schema()
# payload = extractor.run(html=page_html, json_schema=schema)

try:
    product = Product.model_validate(payload)
except ValidationError as e:
    log.warning(e.errors())

The Pydantic documentation covers model_json_schema(), model_validate, and model_dump in full, and it's worth reading before you write your first production model, not after your first outage.

Pro Tip: Run model_json_schema() and print it once during development. If a field's description reads ambiguous to you, it will read ambiguous to the extractor too.

How Do You Write Extraction-Friendly Pydantic Schemas?

Schema design determines extraction quality more than a bigger prompt ever will, according to Pydantic's own best-practice guidance. A handful of habits separate schemas that extract cleanly from ones that generate constant validation noise.

  1. Disambiguate with Field(..., description=...). "Price" is a trap word: list price, sale price, and shipping cost all compete for that label, so spell out which one you mean.
  2. Make fields Optional when they're legitimately missing on some pages. Don't force a str on a field that half your target sites simply don't render.
  3. Nest models for grouped data. A ShippingInfo submodel beats five flat fields prefixed shipping_.
  4. Normalize after extraction, not during. Currency conversion and date parsing belong in a post-processing step, not baked into the schema's job.
  5. Keep every field observable on the page. If a human couldn't point to where a value comes from, don't ask a model to infer it.

Read our structured data extraction with schema validation guide for a longer walkthrough of schema patterns that hold up across hundreds of source domains.

How Should You Fetch and Preprocess Pages Before Extraction?

Fetch strategy matters as much as schema design. A plain requests call is fine for server-rendered pages; anything gated behind client-side JavaScript needs a headless browser like Playwright or a crawler that renders before returning HTML. Our guide to scraping dynamic pages with Playwright covers the tradeoffs in more depth.

Before that HTML ever reaches an extractor, strip it down:

  • Remove <script>, <style>, and navigation boilerplate, since every extra kilobyte is tokens you're paying for and noise the model has to ignore.
  • Tools like html-to-markdown distill a page to its readable content in one call.
  • Use scope selectors (a container div, a main article tag) to point the extractor at the relevant region instead of the whole document.
  • Set sane timeouts, retry with backoff, and send realistic headers. A page that times out silently is worse than one that fails loudly.

What Happens When Extraction or Validation Fails?

Two different failure classes get conflated constantly, and that conflation is where most extraction pipelines rot. A fetch failure means you never got usable HTML. An extraction failure means you got HTML but the returned payload didn't match your schema. Treat them separately.

Structured extraction workflows should check both scrape_status and any json_extraction_error_code the API returns, distinguishing blocked_antibot from a genuine extraction_error versus a plain network failure, per practitioner guidance on typed JSON extraction.

  • Log the raw payload on every failed model_validate call. It's the fastest way to see what the extractor actually saw.
  • Decide per field whether a missing value should retry, fall back to None, or hard-fail the record.
  • Surface structured error metadata (not a bare exception) so downstream consumers can branch on the failure type.

A payload that passes validation still hasn't proven the page was fetched correctly or that the values are current. Add a timestamp and source URL to every record so "valid" and "fresh" don't get confused with each other.

Where Do Validated Pydantic Records Go Next?

A validated model is a checkpoint, not an ending. model_dump() and model_dump_json() are the handoff points into whatever persistence layer you're running, whether that's a Postgres table with matching columns or a document store that accepts nested JSON directly.

  • Persist model_dump() output into schematized tables, or model_dump_json() straight into a document store.
  • Normalize currency values and convert dates to ISO 8601 before anything gets indexed into a vector store, since inconsistent formats quietly break downstream filters.
  • Keep the raw distilled markdown alongside the typed JSON. RAG systems often need the surrounding context that a strict schema necessarily throws away.
  • Record source URL, fetch timestamp, and an extraction confidence signal on every row. Our guide to structured web data for RAG covers this pairing in detail.

The partner guide on data quality checks for prediction market pipelines makes a related point worth borrowing here: passing schema validation is not the same as proving the underlying data was retrieved correctly, and both checks need to run.

What Does LLM-Guided Extraction Cost at Scale?

LLM-assisted extraction costs more per page and runs slower than a cached selector, full stop. That's the tradeoff you're accepting in exchange for not maintaining brittle CSS rules across dozens of layouts.

  • Use a loose json_prompt for discovery while you're still figuring out a site's structure, then freeze the schema once it's stable, a workflow Crawlee's PydanticAiCrawler supports directly.
  • Crawlee's selector-based extractor variant caches CSS selectors once discovered, cutting LLM calls dramatically on repeat runs against the same domain.
  • Batch requests where the extractor supports it. Per-call latency adds up fast at thousand-page scale.

Fixing Common Pydantic Web Extraction Failures

Most extraction problems fall into four buckets, and each has a fast, specific fix rather than a full rewrite.

  1. Extraction returns None for everything. Check scrape_status first. If the fetch itself failed or got blocked, no schema fix will help. Try browser rendering if the site is JavaScript-heavy.
  2. Constant ValidationError on the same fields. Relax those fields to Optional, collect ten failing examples, and look for the actual pattern before touching types again.
  3. Numbers land in the wrong field. Ambiguous Field descriptions are almost always the cause. Refining a description fixes more of these than adjusting the type ever does.
  4. Requests get blocked outright. Respect robots.txt, add exponential backoff, and rotate fetch patterns instead of hammering the same endpoint on a fixed interval.

Pro Tip: When a field keeps extracting wrong values, don't jump to a new model or a bigger prompt. Fix the description first. It's almost always the cheaper, faster win.

Who's Behind This Guidance, and Where to Go Deeper

This guide draws on the same primitives Gyrence exposes as first-class API calls: Fetch for clean, script-stripped HTML and markdown, Extract for schema-guided structured JSON, and Map for discovering a domain's URL graph before you decide what to extract. If you want the longer version of the workflow covered here, our structured JSON extraction guide and our piece on extracting tables from HTML at scale go deeper into production edge cases. The Gyrence docs map each pattern in this article to a specific API call and SDK example.

How Do You Handle Nested Objects and Complex Types?

Real web pages rarely map to flat schemas. A product page has a nested ShippingInfo block, a list of Review objects, and an optional Warranty submodel that only some listings include. Pydantic handles all three natively: nested BaseModel classes, list[Model] for repeated structures, and Optional[Model] for sections that don't always render.

The trap most developers fall into is over-nesting speculatively. If a page has reviews, model reviews: list[Review] = Field(default_factory=list) rather than a fixed-length tuple, because the count varies by page and a fixed length will fail validation the first time a listing has four reviews instead of five. For deeply nested data like a spec table with variable keys, a dict[str, str] field is often more honest than trying to enumerate every possible spec name as a typed field.

Discriminated unions solve a specific, common problem: pages that represent the same entity type in structurally different ways. A marketplace listing might be a PhysicalProduct or a DigitalProduct, each with different required fields. Pydantic's discriminated union support lets the extractor pick the right variant based on a type field, and your downstream code gets a properly typed object instead of a generic dict with optional-everything.

Nested objects and discriminated data variants

One rule holds regardless of nesting depth: every leaf field still needs to trace back to something visible on the page. Nesting organizes complexity; it doesn't excuse inventing structure the source document doesn't actually contain.

How Does Pydantic Fit Into Scrapy or BeautifulSoup?

Pydantic doesn't replace Scrapy or BeautifulSoup. It replaces the untyped dictionary those tools hand you at the end of a parse.

In a Scrapy pipeline, the natural integration point is the Item pipeline stage: parse with Scrapy's selectors as usual, then pass the resulting dict straight into YourModel.model_validate() before yielding it. Validation errors surface immediately, inside the pipeline, instead of three steps downstream when a database insert fails on a null constraint. BeautifulSoup workflows follow the same shape: extract raw values with soup.select() or .find(), assemble them into a dict, validate.

The bigger payoff shows up when you mix selector-based parsing with LLM-assisted fallback. Use BeautifulSoup or Scrapy selectors for the 90% of fields that sit in predictable DOM locations, and route only the genuinely ambiguous fields (a free-text description, an inconsistent price format) to an LLM-guided extractor validated against the same Pydantic model. That hybrid keeps most of your calls cheap and fast while still getting the schema guarantee everywhere.

Crawlee takes this further with dedicated extractor classes: PydanticAiDirectExtractor distills a page directly to a model instance, while PydanticAiSelectorExtractor caches CSS selectors after the first successful extraction to cut repeat LLM costs. Either way, the model you declared stays the single contract every tool in the pipeline has to satisfy.

How Do You Handle Inconsistent or Partial Data?

Partial data isn't an edge case in web extraction. It's the default state of the web. Some listings have a sale price, some don't. Some reviews have a reviewer name, some are anonymous. A schema built as if every field will always be present breaks constantly and for no good reason.

Start by separating "legitimately absent" from "extraction failed." A missing shipping estimate on a digital product is legitimate: model it Optional[str] = None and move on. A missing product name on a product page is a failure worth flagging, because every real product page has one somewhere. Required fields should be reserved for values you'd bet money exist on every page in scope; everything else gets Optional.

For genuinely inconsistent formatting, like a price that shows up as "$19.99", "19.99 USD", or "$19.99*" depending on the page, resist the urge to solve it inside the schema with regex-heavy validators. Accept the raw string, validate it loosely, and normalize in a dedicated post-processing step. That keeps your extraction contract stable even as source formatting drifts.

When partial records are unavoidable, accept them rather than discarding the whole payload. A model_construct() call (Pydantic's validation-skipping constructor) combined with manual field checks lets you persist a partially valid record with clear flags on what's missing, instead of losing 100% of the data because 10% of one field failed.

How Should You Parse Dates, Numbers, and Currency?

Dates and numbers are where most extraction pipelines quietly rot, because the web has no single format for either.

For dates, accept the raw string at extraction time and parse afterward with a dedicated library rather than trying to write a Field type clever enough to catch "March 3, 2026," "03/03/2026," and "3 days ago" all at once. A Pydantic field_validator that calls a robust date-parsing function is more maintainable than a regex pattern baked into the model.

For numbers, currency symbols and thousands separators are the two recurring headaches. "$1,299.00" needs the dollar sign and comma stripped before it becomes a usable float. Pydantic's own best-practice guidance is direct on this: normalize currency and dates after extraction rather than forcing the schema to absorb every formatting variant a source site might use. Keep the raw string in one field and the normalized value in another if you need to audit the conversion later.

For booleans, avoid trusting a model to infer "in stock" from vague page copy. Tie boolean fields to something concrete and observable, like the literal presence of an "Add to Cart" button, and describe that expectation in the Field description so the extractor knows exactly what evidence to look for.

How Do You Keep Pydantic Fast at Scale?

Pydantic v2's validation core is written in Rust, which makes per-record validation fast even at high volume. The bottleneck in large extraction runs is almost never model_validate itself. It's usually the fetch layer, the LLM call, or an oversized schema doing more work than it needs to.

Keep schemas narrow. A model with forty optional fields validates slower and confuses extractors more than five focused models, each scoped to one section of the page. Split a sprawling ProductPage model into Product, ShippingInfo, and Reviews, and only assemble the full record after each piece validates independently.

Batch validation where you can. If you're validating a thousand records from a single crawl, TypeAdapter(list[Product]).validate_python(records) is measurably faster than looping Product.model_validate() a thousand times, because Pydantic's core skips repeated per-call setup overhead.

Cache what doesn't change. If a source's structure has proven stable, freeze the discovery-phase json_prompt into a fixed schema and, where the tooling supports it, a cached selector, cutting both LLM cost and validation surprises on every subsequent run.

Schema-First Extraction Is Becoming the Default, Not a Niche Pattern

Agent-ready pipelines need a contract both sides can trust, and JSON Schema generated from a Pydantic model is turning into that shared language between extraction tools and the systems consuming their output. The practical move is to invest in schema design early, use loose prompt-based discovery while you're still learning a site's shape, then freeze that schema once it earns your trust.

— Glen

Try Gyrence's Extract Primitive on Your Own Pages

Everything in this guide maps directly onto Gyrence's primitives: you declare the JSON Schema, call Extract with that schema against a target URL, and get back a typed response with failure modes surfaced instead of hidden. No separate LLM extraction charge, no surprise line item on the invoice. If a page needs monitoring for drift after you've built the schema, WebDoppler watches it and fires a webhook alert the moment the structure or content changes. Start on the free tier, test your schema against a handful of real URLs, and check the Gyrence docs for the exact request shape Extract expects. When you're ready to run it at volume, current pricing details are available on the Gyrence pricing page, with predictable per-call credit usage to avoid surprise bills.

Sources

FAQ

What Is Pydantic Web Extraction?

Pydantic web extraction is the practice of declaring a typed BaseModel, emitting its JSON Schema, and using that schema to drive a structured or LLM-assisted extractor that returns validated JSON. It gives you a shape guarantee on scraped data, though it doesn't confirm the page was fetched correctly or that the values are current.

Is Pydantic Better Than CSS Selectors for Scraping?

Neither is universally better; they solve different problems. Pydantic-guided extraction wins on unknown or inconsistent markup and multi-site coverage, while selector-based extraction is cheaper and faster once a layout is stable and known.

How Do I Handle a ValidationError During Extraction?

Log the raw payload immediately so you can see what the extractor actually returned, then decide per field whether the fix is relaxing it to Optional, retrying the fetch, or refining the Field description. Most recurring ValidationError cases trace back to an ambiguous description rather than a wrong type, according to practitioner guidance on JSON extraction.

Does Gyrence Support Pydantic-Style Schema Extraction?

Yes. Gyrence's Extract primitive accepts a JSON Schema and returns a typed, validated response with explicit failure modes rather than a silent null, and pricing runs on predictable per-call credits with no separate LLM extraction charge. Current plan pricing is listed on the Gyrence pricing page.

What's the Cost Difference Between LLM Extraction and Selectors?

LLM-guided extraction costs more per page and runs slower than a cached CSS selector, because each call involves a model inference step. A common mitigation is discovering the schema with a loose prompt, then freezing it and caching selectors for repeat runs against the same domain, a pattern Crawlee's extractor variants support directly.