llmcrawl
API reference

Map

Discover every URL on a site without scraping page content.

POST /v1/map

Map combines three sources of URLs, in order:

  1. Sitemaps — robots.txt Sitemap: directives, /sitemap.xml and /sitemap_index.xml, following nested sitemap indexes.
  2. Link discovery — a bounded breadth-first walk of the seed page and its same-site links.
  3. Filtering — the same include/exclude, depth and file-type rules the crawler uses.
FieldTypeDefaultDescription
urlstring—Site to map.
limitnumber5000Maximum links returned (1–5000).
includeSubdomainsbooleantrueInclude subdomains of the base URL.
searchstring—Case-insensitive substring filter on results.
ignoreSitemapbooleantrueSkip sitemap discovery.
includePaths / excludePathsstring[][]Glob filters, as in /v1/crawl.
maxDepthnumber10Maximum path depth.
allowExternalLinksbooleanfalseAllow other domains.
{
  "success": true,
  "links": [
    "https://example.com/pricing",
    "https://example.com/docs",
    "https://example.com/blog/hello"
  ]
}

Map returns URLs only — it does not convert pages to markdown. Use /v1/crawl when you need content, or chain map into batch scrapes when you want to choose the pages yourself.