llmcrawl
API reference

Scrape

Convert a single page into markdown, HTML, links or a screenshot.

POST /v1/scrape

Authenticate with your API key:

Authorization: Bearer llmcrawl_...

Request body

FieldTypeDefaultDescription
urlstring—Page to scrape. A missing scheme becomes https://.
formatsarray["markdown"]Any of markdown, html, rawHtml, links, screenshot, screenshot@fullPage.
headersobject—Extra request headers (cookies, User-Agent, Authorization, …).
includeTagsarray—Keep only these CSS selectors (and their descendants).
excludeTagsarray—Remove these CSS selectors.
onlyMainContentbooleanfalsePrune navigation, footers and sidebars before converting.
timeoutnumber30000Milliseconds, 1000–90000.
waitFornumber0Milliseconds to wait after load before capturing.
extractobject—LLM extraction, see below.
webhookUrlsstring[]—Receive the result by POST as well.
metadataany—Echoed back on the webhook payload.

Unknown fields are rejected rather than ignored, so typos surface immediately.

Response

{
  "success": true,
  "warning": "Page returned status code 404",
  "data": {
    "markdown": "...",
    "links": ["https://example.com/pricing"],
    "screenshot": "iVBORw0KGgo...",
    "extract": { "title": "Example" },
    "metadata": {
      "title": "Example Domain",
      "description": "...",
      "language": "en",
      "ogImage": "https://example.com/og.png",
      "sourceURL": "https://example.com/",
      "statusCode": 200,
      "engine": "playwright"
    }
  }
}

metadata.engine tells you how the page was fetched: remote, playwright or fetch. See JavaScript rendering.

Error status codes

StatusMeaning
400Invalid body. details describes what was wrong.
401Missing or invalid API key.
429Rate limit exceeded for this endpoint.
502Every way of fetching the page failed.

Structured extraction

Provide a JSON Schema and llmcrawl will return data matching it:

curl -X POST "$LLMCRAWL_API_URL/v1/scrape" \
  -H "Authorization: Bearer $LLMCRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/product",
    "formats": ["markdown"],
    "extract": {
      "schema": {
        "type": "object",
        "properties": {
          "name": { "type": "string" },
          "price": { "type": "number" }
        },
        "required": ["name", "price"]
      }
    }
  }'

If extraction is not available on the deployment you are calling, the request still succeeds and metadata.extractSkipped is true.

Keeping URLs out of your own network

Every URL is checked before it is fetched. Loopback, private, link-local and cloud metadata addresses are rejected, and hostnames are re-checked after DNS resolution so a rebinding host cannot slip through. Self-hosters who need to scrape internal targets can opt in with ALLOW_PRIVATE_ADDRESSES — see Self-hosting.

On this page