Structured extraction
Return JSON that matches your schema instead of scraping fields yourself.
Structured extraction turns a page into the JSON your application wants. You describe the
data with a JSON Schema, add an extract block to a /v1/scrape request, and the page comes
back with a matching data.extract object — no parsing or selector maintenance on your side.
Extract a page
curl -X POST "$LLMCRAWL_API_URL/v1/scrape" \
-H "Authorization: Bearer $LLMCRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/product",
"formats": ["markdown"],
"extract": {
"schema": {
"type": "object",
"properties": {
"name": { "type": "string", "description": "Product name" },
"price": { "type": "number", "description": "Current price in USD" },
"in_stock": { "type": "boolean" }
},
"required": ["name", "price"]
}
}
}'{
"success": true,
"data": {
"markdown": "...",
"extract": {
"name": "Example Wireless Headphones",
"price": 129.99,
"in_stock": true
},
"metadata": { "sourceURL": "https://example.com/product", "statusCode": 200 }
}
}Field descriptions matter: describe what each field means and where it lives on the page,
and the extraction gets noticeably more reliable.
Extract options
| Field | Type | Default | Description |
|---|---|---|---|
schema | object | — | JSON Schema describing the data to extract. |
mode | string | "llm" | Extraction mode. |
systemPrompt | string | built-in | Overrides the instruction sent to the model. |
prompt | string | — | Extra page-specific guidance for this request. |
When extraction is unavailable
If extraction is not available for your request, the scrape still succeeds: data.extract is
absent, a human-readable warning is returned, and metadata.extractSkipped is true. Treat
it as a soft failure and fall back to markdown parsing.
Build schemas in the dashboard
The dashboard has a schema builder under Extraction: start from a template (product, article, company, event, restaurant) or build a custom schema field by field, add descriptions, instructions and validation rules, and run it against a live page before you automate it.
Saved schemas are reusable — copy the generated JSON Schema into the extract block of your
API calls so what you tested in the dashboard is exactly what runs in production.