llmcrawl

Introduction

Turn websites into LLM-ready data — markdown, HTML, links, screenshots and structured JSON.

llmcrawl turns web pages into clean, structured data your applications and agents can use: markdown, HTML, links, screenshots and schema-driven JSON.

Point it at a single page or a whole site — from the dashboard or the REST API — and get content you can index, fine-tune on, feed to an agent or store in your own pipeline.

What you can do

  • Scrape — one page in one request: markdown, HTML, raw HTML, links and screenshots.
  • Crawl — follow a whole site with include/exclude path filters, depth limits and robots.txt support.
  • Map — discover the URLs on a site before deciding what to scrape.
  • Extract — return structured JSON matching a schema you define.
  • Webhooks — receive a POST when a scrape finishes or a crawl completes.

How it works

  1. Call the API with your API key — or pay per request with x402 and no account at all.
  2. llmcrawl fetches the page, rendering JavaScript in a real browser when needed, and retries with a browser-grade request if a site resists the first attempt.
  3. The page is converted to the formats you asked for and returned as JSON.

The API base URL

Every endpoint lives under a single base URL — the hosted API at https://api.llmcrawl.dev. Examples in these docs use:

export LLMCRAWL_API_URL=https://api.llmcrawl.dev
export LLMCRAWL_API_KEY=llmcrawl_your_key_here

On this page