llmcrawl

Introduction

Turn websites into LLM-ready data — markdown, HTML, links, screenshots and structured JSON.

llmcrawl fetches web pages and converts them into clean, structured data that large language models can consume: markdown, HTML, links, screenshots and schema-driven JSON.

Point it at a page or a whole site and you get content you can index, fine-tune on, feed to an agent or store in your own pipeline.

What you get

  • Scrape — one page, in one request. Markdown, HTML, raw HTML, links, screenshots.
  • Crawl — follow a whole site with include/exclude path filters, depth limits and robots.txt support.
  • Map — discover every URL on a site from its sitemap and by walking links, without fetching page content.
  • Extract — pull structured JSON from a page using a JSON Schema you define.

How it works

  1. You call the API with your API key — or pay per request with x402 and no account at all.
  2. llmcrawl fetches the page, rendering JavaScript in a real browser when needed, and escalates automatically if a site resists the first attempt.
  3. The page is converted to the formats you asked for and returned as JSON.

The API base URL

Every endpoint lives under a single base URL — for example https://api.llmcrawl.dev for the hosted API, or the URL of your own self-hosted deployment. Examples in these docs use:

export LLMCRAWL_API_URL=https://api.llmcrawl.dev
export LLMCRAWL_API_KEY=llmcrawl_your_key_here

On this page