JavaScript rendering
How llmcrawl fetches pages, renders client-side content, and captures screenshots.
Not every page is a static HTML file. llmcrawl picks a fetching strategy per request: plain HTTP for static pages, a real headless browser for anything that needs JavaScript, and automatic escalation when a site resists the first attempt.
You do not have to choose — ask for what you need and llmcrawl routes the request to a strategy that can satisfy it.
What the engine field means
Every response includes metadata.engine, which tells you how the page was fetched:
| Engine | When it is used |
|---|---|
fetch | The page was fetched over plain HTTP. Fastest path, right for static pages. |
playwright | The page was rendered in a headless browser, so client-side JavaScript ran. |
remote | The page was rendered by a remote browser pool. Used when the request needs a browser and the rendering is delegated. |
A request that needs a browser (for example a screenshot) is never silently downgraded to a plain fetch — it either succeeds with a rendered page or fails.
Waiting for client-side content
If a page loads its content after the initial HTML arrives, use waitFor to pause after load before capturing:
curl -X POST "$LLMCRAWL_API_URL/v1/scrape" \
-H "Authorization: Bearer $LLMCRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/app", "formats": ["markdown"], "waitFor": 3000}'waitFor is in milliseconds and can only be satisfied by a browser-rendered request. Use it sparingly — it adds the wait to every request.
Screenshots
Add screenshot to formats for a viewport-sized capture, or screenshot@fullPage for the entire scrollable page. Screenshots come back as base64-encoded PNG in data.screenshot:
curl -X POST "$LLMCRAWL_API_URL/v1/scrape" \
-H "Authorization: Bearer $LLMCRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com", "formats": ["screenshot@fullPage"]}'Images, media and fonts are blocked while rendering to keep captures fast, so text and layout are faithful but embedded images are not loaded.
Content selection
Three options shape what ends up in the markdown:
| Option | Effect |
|---|---|
onlyMainContent | Prunes navigation, footers and sidebars before conversion. |
includeTags | Keeps only the elements matching these CSS selectors. |
excludeTags | Removes elements matching these CSS selectors. |
They also apply to crawls via scrapeOptions.
Sites that block bots
Sites behind bot protection or rate limiters are handled automatically: when a page is detected as a challenge, llmcrawl re-attempts it with a browser-grade request fingerprint. From your side this is transparent — the request either returns the rendered page or returns 502 (see Scrape) if every strategy failed.
Authenticated pages
Pass cookies or authorization headers with the headers field:
curl -X POST "$LLMCRAWL_API_URL/v1/scrape" \
-H "Authorization: Bearer $LLMCRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://app.example.com/private", "formats": ["markdown"], "headers": {"Cookie": "session=..."}}'Timeouts
timeout bounds the whole request in milliseconds (1000–90000, default 30000). Timeouts apply to fetching and rendering combined, so a browser-rendered screenshot of a heavy page can need a longer timeout than a markdown scrape of the same page.