API reference
Crawl
Follow a site across many pages and read the results back.
Crawling is asynchronous: POST /v1/crawl returns an id immediately, llmcrawl processes the site in the background, and you poll GET /v1/crawl/{id} for progress and documents.
Start a crawl
POST /v1/crawl
| Field | Type | Default | Description |
|---|---|---|---|
url | string | — | Seed URL. |
limit | number | 10000 | Maximum pages to scrape. |
maxDepth | number | 10 | Maximum path depth below the seed. |
maxDiscoveryDepth | number | — | Maximum number of link hops from the seed. |
includePaths | string[] | [] | Glob patterns a URL must match. Empty means everything. |
excludePaths | string[] | [] | Glob patterns that reject a URL. |
allowBackwardLinks | boolean | false | Allow URLs above the seed path. |
allowExternalLinks | boolean | false | Allow other domains. |
ignoreSitemap | boolean | true | Skip sitemap seeding. |
ignoreRobotsTxt | boolean | false | Ignore robots.txt rules. |
regexOnFullURL | boolean | false | Match patterns against the full URL instead of the path. |
scrapeOptions | object | — | Same options as /v1/scrape, minus timeout. |
webhookUrls | string[] | — | Receive a crawl.completed event. |
webhookMetadata | any | — | Attached to webhook payloads. |
Response:
{ "success": true, "id": "crawl_9f2c...", "url": "https://example.com" }Check progress
GET /v1/crawl/{id}?offset=0&limit=100
{
"success": true,
"status": "scraping",
"completed": 12,
"total": 34,
"expiresAt": "2026-09-22T10:15:00.000Z",
"next": "/v1/crawl/crawl_9f2c...?offset=100",
"data": [{ "markdown": "...", "metadata": { "sourceURL": "https://example.com/about" } }]
}status is one of scraping, completed, failed or cancelled. next is present while more documents are available.
Cancel a crawl
DELETE /v1/crawl/{id} marks the crawl cancelled, removes jobs that have not started, and makes running jobs stop before their next page.
What a crawl will and won't follow
- Each discovered URL is normalised before queuing, so
http://www.example.com/a/andhttps://example.com/acount as one page and are never scraped twice. - A crawl is
completedonce link discovery from the seed has finished and every queued page has been processed. - Skipped URLs are counted by reason (
exclude_match,depth_limit,external_domain, …), which makes it easy to tell whether a filter or a limit is holding your crawl back.
Crawled documents are available for 24 hours (CRAWL_TTL_SECONDS on self-hosted deployments); store anything you need to keep.