llmcrawl

Structured extraction

Return JSON that matches your schema instead of scraping fields yourself.

Structured extraction turns a page into the JSON your application wants. You describe the data with a JSON Schema, add an extract block to a /v1/scrape request, and the page comes back with a matching data.extract object — no parsing or selector maintenance on your side.

Extract a page

curl -X POST "$LLMCRAWL_API_URL/v1/scrape" \
  -H "Authorization: Bearer $LLMCRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/product",
    "formats": ["markdown"],
    "extract": {
      "schema": {
        "type": "object",
        "properties": {
          "name": { "type": "string", "description": "Product name" },
          "price": { "type": "number", "description": "Current price in USD" },
          "in_stock": { "type": "boolean" }
        },
        "required": ["name", "price"]
      }
    }
  }'
{
  "success": true,
  "data": {
    "markdown": "...",
    "extract": {
      "name": "Example Wireless Headphones",
      "price": 129.99,
      "in_stock": true
    },
    "metadata": { "sourceURL": "https://example.com/product", "statusCode": 200 }
  }
}

Field descriptions matter: describe what each field means and where it lives on the page, and the extraction gets noticeably more reliable.

Extract options

FieldTypeDefaultDescription
schemaobject—JSON Schema describing the data to extract.
modestring"llm"Extraction mode.
systemPromptstringbuilt-inOverrides the instruction sent to the model.
promptstring—Extra page-specific guidance for this request.

When extraction is unavailable

If extraction is not available for your request, the scrape still succeeds: data.extract is absent, a human-readable warning is returned, and metadata.extractSkipped is true. Treat it as a soft failure and fall back to markdown parsing.

Build schemas in the dashboard

The dashboard has a schema builder under Extraction: start from a template (product, article, company, event, restaurant) or build a custom schema field by field, add descriptions, instructions and validation rules, and run it against a live page before you automate it.

Saved schemas are reusable — copy the generated JSON Schema into the extract block of your API calls so what you tested in the dashboard is exactly what runs in production.

On this page