llmcrawl

Quickstart

Get an API key and make your first scrape, crawl and map.

1. Get an API key

Sign in to the llmcrawl dashboard, open API keys and create a key. The full key is shown once — copy it somewhere safe.

If you would rather not create an account, you can pay per request with x402 instead of using a key.

2. Set your environment

export LLMCRAWL_API_URL=https://api.llmcrawl.dev
export LLMCRAWL_API_KEY=llmcrawl_your_key_here

LLMCRAWL_API_URL is the base URL of the API you signed up for, or of your own deployment if you are self-hosting.

3. Scrape a page

curl -X POST "$LLMCRAWL_API_URL/v1/scrape" \
  -H "Authorization: Bearer $LLMCRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com", "formats": ["markdown"]}'
{
  "success": true,
  "data": {
    "markdown": "# Example Domain\n\nThis domain is for use in illustrative examples...",
    "metadata": {
      "title": "Example Domain",
      "sourceURL": "https://example.com/",
      "statusCode": 200,
      "engine": "playwright"
    }
  }
}

Every call takes the same shape: authenticate, pass a JSON body, get a JSON response. See Scrape for every field and format.

4. Map a site

Before crawling, see what URLs exist:

curl -X POST "$LLMCRAWL_API_URL/v1/map" \
  -H "Authorization: Bearer $LLMCRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com", "limit": 100}'
{
  "success": true,
  "links": ["https://example.com/", "https://example.com/about"]
}

Map returns URLs only and does not fetch page content.

5. Crawl a site

curl -X POST "$LLMCRAWL_API_URL/v1/crawl" \
  -H "Authorization: Bearer $LLMCRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com", "limit": 25, "maxDepth": 2}'

Crawling is asynchronous: the response returns an id immediately, pages are processed in the background, and you poll until the status is completed:

curl "$LLMCRAWL_API_URL/v1/crawl/$CRAWL_ID" \
  -H "Authorization: Bearer $LLMCRAWL_API_KEY"
{
  "success": true,
  "status": "scraping",
  "completed": 12,
  "total": 34,
  "data": [{ "markdown": "...", "metadata": { "sourceURL": "https://example.com/about" } }]
}

6. Extract structured JSON

Add an extract block with a JSON Schema and the page is returned as matching JSON:

curl -X POST "$LLMCRAWL_API_URL/v1/scrape" \
  -H "Authorization: Bearer $LLMCRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/product",
    "formats": ["markdown"],
    "extract": {
      "schema": {
        "type": "object",
        "properties": { "name": { "type": "string" }, "price": { "type": "number" } },
        "required": ["name", "price"]
      }
    }
  }'

From JavaScript or Python

The API is plain HTTP, so any HTTP client works. In Node.js:

const response = await fetch(`${process.env.LLMCRAWL_API_URL}/v1/scrape`, {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LLMCRAWL_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ url: "https://example.com", formats: ["markdown"] }),
});

const { data } = await response.json();
console.log(data.markdown);

Next steps

On this page