Introduction
Turn websites into LLM-ready data — markdown, HTML, links, screenshots and structured JSON.
llmcrawl fetches web pages and converts them into clean, structured data that large language models can consume: markdown, HTML, links, screenshots and schema-driven JSON.
Point it at a page or a whole site and you get content you can index, fine-tune on, feed to an agent or store in your own pipeline.
What you get
- Scrape — one page, in one request. Markdown, HTML, raw HTML, links, screenshots.
- Crawl — follow a whole site with include/exclude path filters, depth limits and robots.txt support.
- Map — discover every URL on a site from its sitemap and by walking links, without fetching page content.
- Extract — pull structured JSON from a page using a JSON Schema you define.
How it works
- You call the API with your API key — or pay per request with x402 and no account at all.
- llmcrawl fetches the page, rendering JavaScript in a real browser when needed, and escalates automatically if a site resists the first attempt.
- The page is converted to the formats you asked for and returned as JSON.
The API base URL
Every endpoint lives under a single base URL — for example https://api.llmcrawl.dev for the hosted API, or the URL of your own self-hosted deployment. Examples in these docs use:
export LLMCRAWL_API_URL=https://api.llmcrawl.dev
export LLMCRAWL_API_KEY=llmcrawl_your_key_here