llmcrawl

Self-hosting

Run llmcrawl on your own infrastructure — Docker Compose, environment variables and scaling.

This page is for running your own llmcrawl deployment. If you are using the hosted API, you only need the Quickstart and the API reference.

Requirements

  • Bun 1.4+ (or Docker)
  • PostgreSQL 18+
  • Redis 7+

Docker Compose

The quickest path is the included docker-compose.yml, which builds the api, worker and browser services and runs Postgres and Redis alongside them:

docker compose up -d --build
docker compose logs -f worker

Secrets are mounted as BuildKit secrets so nothing sensitive lands in an image layer.

Only the browser image installs Chromium. The API and worker images get their rendering from BROWSER_ENGINE_URL=http://browser:3002 and never need a browser locally. The browser service is not published to the host — it is reachable only on the Compose network. If you expose it, set BROWSER_ENGINE_TOKEN and the matching value on the API and worker.

Environment variables

A missing or malformed required variable fails at boot with a message naming it, so start with the required columns below and add the rest as needed. Environment files (.env, .env.<NODE_ENV>, .env.local) are read at process start — restart the service after editing one.

Shared

VariableRequiredDescription
DATABASE_URLyesPostgres connection string.
REDIS_URLyesRedis connection string.
LLMCRAWL_USER_AGENTnoUser agent sent to targets.
ALLOW_PRIVATE_ADDRESSESnoAllow loopback/private targets. Development only.
BROWSER_ENGINE_URLnoDelegate rendering to a browser service.
BROWSER_ENGINE_TOKENnoShared bearer token for that service.
LLM_API_KEY / LLM_BASE_URL / LLM_MODELnoOpenAI-compatible endpoint used for structured extraction. Without it, extraction is skipped and scrape responses carry metadata.extractSkipped.

API

VariableRequiredDescription
BETTER_AUTH_SECRETyes32+ characters.
BETTER_AUTH_URLyesPublic URL of this API.
CORS_ORIGINyesComma-separated allowed origins.
API_KEY_PEPPERyes16+ characters. Mixed into API key hashes.
CRAWL_TTL_SECONDSnoRedis TTL for crawl state (default 86400).
SCRAPE_TIMEOUT_MSnoDefault scrape timeout.
POLAR_ACCESS_TOKEN / POLAR_SUCCESS_URLnoSubscription billing, when enabled.
X402_ENABLED / X402_PAY_TOnoMachine payments. Setting a receiving address enables x402.
X402_FACILITATOR_URLnoFacilitator URL. Development defaults to the public testnet service.
X402_NETWORKnoeip155:84532 in development, eip155:8453 in production.
X402_PRICEnoDefault per-operation price ($0.01); endpoint-specific overrides are available.

Worker and browser

VariableDefaultDescription
WORKER_CONCURRENCY2Pages rendered in parallel per worker.
PORT3002Browser service port.
PLAYWRIGHT_MAX_CONTEXTS4Browser contexts per process.
BROWSER_MAX_USES_PER_CONTEXT200Recycle a context after this many pages.
BROWSER_IDLE_TIMEOUT_MS60000Close idle contexts after this long.

Payments (x402)

x402 is disabled by default. Setting a receiving address enables it. For Base Sepolia development:

X402_ENABLED=true
X402_PAY_TO=0xYourReceivingWallet
X402_NETWORK=eip155:84532
X402_FACILITATOR_URL=https://x402.org/facilitator
X402_PRICE='$0.01'

In production, configure a production facilitator explicitly. The default network is Base mainnet and the testnet-only x402.org facilitator is rejected when NODE_ENV=production:

NODE_ENV=production
X402_ENABLED=true
X402_PAY_TO=0xYourBaseWallet
X402_NETWORK=eip155:8453
X402_FACILITATOR_URL=https://your-production-facilitator.example
X402_PRICE='$0.01'

Every operation is priced independently; endpoint-specific values override X402_PRICE:

VariableProtected operation
X402_SCRAPE_PRICEPOST /v1/scrape
X402_CRAWL_PRICEPOST /v1/crawl
X402_CRAWL_STATUS_PRICEGET /v1/crawl/:id
X402_CRAWL_CANCEL_PRICEDELETE /v1/crawl/:id
X402_MAP_PRICEPOST /v1/map
VariableDefaultDescription
X402_FACILITATOR_BEARER_TOKEN—Bearer credential for a managed or self-hosted facilitator.
X402_FACILITATOR_TIMEOUT_MS30000Timeout for facilitator verify, settle, and capability requests.
X402_MAX_TIMEOUT_SECONDS60Payment authorization validity window, up to 300 seconds.

Scaling

  • Workers. Add replicas. Jobs are distributed across workers and crawls stay deduplicated.
  • Browsers. Add replicas behind a load balancer and point BROWSER_ENGINE_URL at it. The service is stateless.
  • API. Add replicas behind a load balancer. Rate limits are shared, with an in-memory fallback if Redis is briefly unavailable.
  • Postgres. Indexed per user, per crawl and per endpoint so the dashboard and status polling stay fast.

Operations

  • GET /health — liveness.
  • GET /ready — readiness; checks Postgres and Redis and returns 503 when either is down.
  • GET /api-reference — generated OpenAPI reference.

Running behind a proxy

Set BETTER_AUTH_URL and CORS_ORIGIN to your public URLs. Cookies are issued with sameSite: "none" and secure: true, so TLS is required in production.

On this page