llmcrawl
API reference

Crawl

Follow a site across many pages and read the results back.

Crawling is asynchronous: POST /v1/crawl returns an id immediately, llmcrawl processes the site in the background, and you poll GET /v1/crawl/{id} for progress and documents.

Start a crawl

POST /v1/crawl

FieldTypeDefaultDescription
urlstring—Seed URL.
limitnumber10000Maximum pages to scrape.
maxDepthnumber10Maximum path depth below the seed.
maxDiscoveryDepthnumber—Maximum number of link hops from the seed.
includePathsstring[][]Glob patterns a URL must match. Empty means everything.
excludePathsstring[][]Glob patterns that reject a URL.
allowBackwardLinksbooleanfalseAllow URLs above the seed path.
allowExternalLinksbooleanfalseAllow other domains.
ignoreSitemapbooleantrueSkip sitemap seeding.
ignoreRobotsTxtbooleanfalseIgnore robots.txt rules.
regexOnFullURLbooleanfalseMatch patterns against the full URL instead of the path.
scrapeOptionsobject—Same options as /v1/scrape, minus timeout.
webhookUrlsstring[]—Receive a crawl.completed event.
webhookMetadataany—Attached to webhook payloads.

Response:

{ "success": true, "id": "crawl_9f2c...", "url": "https://example.com" }

Check progress

GET /v1/crawl/{id}?offset=0&limit=100

{
  "success": true,
  "status": "scraping",
  "completed": 12,
  "total": 34,
  "expiresAt": "2026-09-22T10:15:00.000Z",
  "next": "/v1/crawl/crawl_9f2c...?offset=100",
  "data": [{ "markdown": "...", "metadata": { "sourceURL": "https://example.com/about" } }]
}

status is one of scraping, completed, failed or cancelled. next is present while more documents are available.

Cancel a crawl

DELETE /v1/crawl/{id} marks the crawl cancelled, removes jobs that have not started, and makes running jobs stop before their next page.

What a crawl will and won't follow

  • Each discovered URL is normalised before queuing, so http://www.example.com/a/ and https://example.com/a count as one page and are never scraped twice.
  • A crawl is completed once link discovery from the seed has finished and every queued page has been processed.
  • Skipped URLs are counted by reason (exclude_match, depth_limit, external_domain, …), which makes it easy to tell whether a filter or a limit is holding your crawl back.

Crawled documents are available for 24 hours (CRAWL_TTL_SECONDS on self-hosted deployments); store anything you need to keep.

On this page