Introduction
Turn websites into LLM-ready data — markdown, HTML, links, screenshots and structured JSON.
llmcrawl turns web pages into clean, structured data your applications and agents can use: markdown, HTML, links, screenshots and schema-driven JSON.
Point it at a single page or a whole site — from the dashboard or the REST API — and get content you can index, fine-tune on, feed to an agent or store in your own pipeline.
What you can do
- Scrape — one page in one request: markdown, HTML, raw HTML, links and screenshots.
- Crawl — follow a whole site with include/exclude path filters, depth limits and robots.txt support.
- Map — discover the URLs on a site before deciding what to scrape.
- Extract — return structured JSON matching a schema you define.
- Webhooks — receive a POST when a scrape finishes or a crawl completes.
How it works
- Call the API with your API key — or pay per request with x402 and no account at all.
- llmcrawl fetches the page, rendering JavaScript in a real browser when needed, and retries with a browser-grade request if a site resists the first attempt.
- The page is converted to the formats you asked for and returned as JSON.
The API base URL
Every endpoint lives under a single base URL — the hosted API at https://api.llmcrawl.dev.
Examples in these docs use:
export LLMCRAWL_API_URL=https://api.llmcrawl.dev
export LLMCRAWL_API_KEY=llmcrawl_your_key_here