Site map (URL discovery)

$0.005 per call · USDC via x402 · POST /api/site-map

Discover a website's URLs in one call: reads robots.txt, its declared sitemap(s) (sitemap indexes and gzipped sitemaps included, /sitemap.xml as the fallback) and the start page's internal links, then returns a same-host, normalized, deduplicated list (up to 500) with an optional substring filter. Send POST /api/site-map with the required field url and pay $0.005 per call over x402 or MPP (there is no free tier). It returns a JSON object with url, host, total, urls, sources and 7 more.

Hard budgets: at most 6 fetches, 15 seconds, 5 MB. Use it to pick which pages to crawl or extract next.

Category: Web & documents · Tags: web sitemap crawl urls discovery seo robots site

TRY IN PLAYGROUND →

Parameters

NameTypeRequiredDescription
urlstringyesStart URL (the site's homepage or any page on it) Also accepted as link, uri, href, page.
limitintegernoMax URLs to return, 1-500 (default 100)
includeSubdomainsbooleannoAlso keep URLs on subdomains of the start site (default false; www and bare host always count as one site)
searchstringnoOptional case-insensitive substring filter applied to the discovered URLs

Example request

curl -i -X POST https://agent402.tools/api/site-map \
  -H "Content-Type: application/json" \
  -d '{"url":"https://www.iana.org","limit":50}'

Without payment this returns HTTP 402 Payment Required with the exact price for site-map; any x402 v2 or MPP client pays it and retries.

Example response

{
  "url": "https://www.iana.org/",
  "host": "www.iana.org",
  "total": 120,
  "urls": [
    "https://www.iana.org/",
    "https://www.iana.org/domains",
    "https://www.iana.org/numbers",
    "https://www.iana.org/protocols"
  ],
  "sources": {
    "sitemap": 96,
    "links": 24
  },
  "sitemapsRead": 1,
  "truncated": true,
  "search": null,
  "fetches": 3,
  "warnings": [],
  "source": "robots.txt, sitemap(s) and start-page links, fetched live",
  "fetchedAt": "2026-08-22T00:00:00.000Z"
}
FieldTypeAlways presentIn the example
urlstringyeshttps://www.iana.org/
hoststringyeswww.iana.org
totalnumberyes120
urlsarray of stringyes4 items in the example
sourcesobjectyes2 fields: sitemap, links
sitemapsReadnumberyes1
truncatedbooleanyestrue
searchnullnonull
fetchesnumberyes3
warningsarrayyes0 items in the example
sourcestringyesrobots.txt, sitemap(s) and start-page links, fetched live
fetchedAtstringyes2026-08-22T00:00:00.000Z

From an MCP client

catalog.call {
  "slug": "site-map",
  "params": {
    "url": "https://www.iana.org",
    "limit": 50
  }
}

The hosted connector at https://agent402.tools/mcp needs a payment for site-map; the stdio package pays it from a wallet or from AGENT402_CREDITS_KEY. Local install: npx -y agent402-mcp.

Errors and behavior

Paid call (JavaScript agent)

import { wrapFetchWithPayment } from "@x402/fetch";
import { x402Client } from "@x402/core/client";
import { registerExactEvmScheme } from "@x402/evm/exact/client";
import { privateKeyToAccount } from "viem/accounts";

const client = new x402Client();
client.setSpendControls?.(false); // keep your own spending ceiling in code
registerExactEvmScheme(client, { signer: privateKeyToAccount(KEY) });
const payFetch = wrapFetchWithPayment(fetch, client);

const res = await payFetch("https://agent402.tools/api/site-map", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    "url": "https://www.iana.org",
    "limit": 50
  }),
});

Related tools

Site crawl (pages to markdown)

$0.02 · POST /api/site-crawl

Crawl a website breadth-first from a start URL over its internal links and return each page as clean markdown (or plain …

Sitemap reader

$0.002 · POST /api/sitemap

Fetch and parse a sitemap.xml (or sitemap index): returns up to 500 URLs with lastmod, or the child sitemaps of an index…

Wayback Machine snapshot

$0.003 · POST /api/archive-snapshot

Look up a URL in the Internet Archive's Wayback Machine: returns the archived snapshot closest to an optional timestamp …

Exa grounded answer

$0.010 · POST /api/exa-answer

Ask a question and get a written answer with the sources it was drawn from. Exa searches its index, reads the pages and …

Exa page contents

$0.006 · POST /api/exa-contents

Retrieve the readable text of up to 10 web pages by URL, with optional query-focused highlights. Exa serves from its cra…

Exa neural web search

$0.012 · POST /api/exa-search

Neural (embedding-based) web search over Exa's curated index, which finds pages by meaning rather than keyword overlap. …