Scrape
Convert a single page into markdown, HTML, links or a screenshot.
POST /v1/scrape
Authenticate with your API key:
Authorization: Bearer llmcrawl_...Request body
| Field | Type | Default | Description |
|---|---|---|---|
url | string | — | Page to scrape. A missing scheme becomes https://. |
formats | array | ["markdown"] | Any of markdown, html, rawHtml, links, screenshot, screenshot@fullPage. |
headers | object | — | Extra request headers (cookies, User-Agent, Authorization, …). |
includeTags | array | — | Keep only these CSS selectors (and their descendants). |
excludeTags | array | — | Remove these CSS selectors. |
onlyMainContent | boolean | false | Prune navigation, footers and sidebars before converting. |
timeout | number | 30000 | Milliseconds, 1000–90000. |
waitFor | number | 0 | Milliseconds to wait after load before capturing. |
extract | object | — | LLM extraction, see below. |
webhookUrls | string[] | — | Receive the result by POST as well. |
metadata | any | — | Echoed back on the webhook payload. |
Unknown fields are rejected rather than ignored, so typos surface immediately.
Response
{
"success": true,
"warning": "Page returned status code 404",
"data": {
"markdown": "...",
"links": ["https://example.com/pricing"],
"screenshot": "iVBORw0KGgo...",
"extract": { "title": "Example" },
"metadata": {
"title": "Example Domain",
"description": "...",
"language": "en",
"ogImage": "https://example.com/og.png",
"sourceURL": "https://example.com/",
"statusCode": 200,
"engine": "playwright"
}
}
}metadata.engine tells you how the page was fetched: remote, playwright or fetch. See JavaScript rendering.
Error status codes
| Status | Meaning |
|---|---|
400 | Invalid body. details describes what was wrong. |
401 | Missing or invalid API key. |
429 | Rate limit exceeded for this endpoint. |
502 | Every way of fetching the page failed. |
Structured extraction
Provide a JSON Schema and llmcrawl will return data matching it:
curl -X POST "$LLMCRAWL_API_URL/v1/scrape" \
-H "Authorization: Bearer $LLMCRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/product",
"formats": ["markdown"],
"extract": {
"schema": {
"type": "object",
"properties": {
"name": { "type": "string" },
"price": { "type": "number" }
},
"required": ["name", "price"]
}
}
}'If extraction is not available on the deployment you are calling, the request still succeeds and metadata.extractSkipped is true.
Keeping URLs out of your own network
Every URL is checked before it is fetched. Loopback, private, link-local and cloud metadata addresses are rejected, and hostnames are re-checked after DNS resolution so a rebinding host cannot slip through. Self-hosters who need to scrape internal targets can opt in with ALLOW_PRIVATE_ADDRESSES — see Self-hosting.