llmcrawl

JavaScript rendering

How llmcrawl fetches pages, renders client-side content, and captures screenshots.

Not every page is a static HTML file. llmcrawl picks a fetching strategy per request: plain HTTP for static pages, a real headless browser for anything that needs JavaScript, and automatic escalation when a site resists the first attempt.

You do not have to choose — ask for what you need and llmcrawl routes the request to a strategy that can satisfy it.

What the engine field means

Every response includes metadata.engine, which tells you how the page was fetched:

EngineWhen it is used
fetchThe page was fetched over plain HTTP. Fastest path, right for static pages.
playwrightThe page was rendered in a headless browser, so client-side JavaScript ran.
remoteThe page was rendered by a remote browser pool. Used when the request needs a browser and the rendering is delegated.

A request that needs a browser (for example a screenshot) is never silently downgraded to a plain fetch — it either succeeds with a rendered page or fails.

Waiting for client-side content

If a page loads its content after the initial HTML arrives, use waitFor to pause after load before capturing:

curl -X POST "$LLMCRAWL_API_URL/v1/scrape" \
  -H "Authorization: Bearer $LLMCRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/app", "formats": ["markdown"], "waitFor": 3000}'

waitFor is in milliseconds and can only be satisfied by a browser-rendered request. Use it sparingly — it adds the wait to every request.

Screenshots

Add screenshot to formats for a viewport-sized capture, or screenshot@fullPage for the entire scrollable page. Screenshots come back as base64-encoded PNG in data.screenshot:

curl -X POST "$LLMCRAWL_API_URL/v1/scrape" \
  -H "Authorization: Bearer $LLMCRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com", "formats": ["screenshot@fullPage"]}'

Images, media and fonts are blocked while rendering to keep captures fast, so text and layout are faithful but embedded images are not loaded.

Content selection

Three options shape what ends up in the markdown:

OptionEffect
onlyMainContentPrunes navigation, footers and sidebars before conversion.
includeTagsKeeps only the elements matching these CSS selectors.
excludeTagsRemoves elements matching these CSS selectors.

They also apply to crawls via scrapeOptions.

Sites that block bots

Sites behind bot protection or rate limiters are handled automatically: when a page is detected as a challenge, llmcrawl re-attempts it with a browser-grade request fingerprint. From your side this is transparent — the request either returns the rendered page or returns 502 (see Scrape) if every strategy failed.

Authenticated pages

Pass cookies or authorization headers with the headers field:

curl -X POST "$LLMCRAWL_API_URL/v1/scrape" \
  -H "Authorization: Bearer $LLMCRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://app.example.com/private", "formats": ["markdown"], "headers": {"Cookie": "session=..."}}'

Timeouts

timeout bounds the whole request in milliseconds (1000–90000, default 30000). Timeouts apply to fetching and rendering combined, so a browser-rendered screenshot of a heavy page can need a longer timeout than a markdown scrape of the same page.

On this page