Crawl a docs site
What the crawler reads, how include and exclude paths work, and how to keep an index fresh.
The crawler is one of three ways in — it reads sitemaps and server-rendered HTML, and turns pages into hierarchy records. It does not run JavaScript, so a page's searchable content must be present in the server-rendered HTML.
Starting a crawl
Paste your URL in the dashboard, or run:
npx @suo/cli crawl --url https://docs.example.comsuo looks for sitemap.xml at your domain root first. If it finds one, every URL listed is a candidate page. If it does not, the crawler follows links starting from the URL you gave it, staying on the same domain.
Include and exclude paths
Adjust which paths are crawled with chips in the dashboard, or in suo.config.ts:
export default defineConfig({
index: 'docs',
source: {
type: 'crawl',
url: 'https://docs.example.com',
include: ['/docs/**'],
exclude: ['/docs/internal/**', '/docs/*/changelog'],
},
});A path excluded here is skipped before a request is even made.
What becomes a record
For each page, the crawler reads the <title>, the heading structure (h1–h6) and the body text in between, and produces one record per heading section, each carrying its place in the page's hierarchy (lvl0–lvl6). Navigation chrome, footers and anything marked data-suo-ignore are skipped.
<main data-suo-content>
<h1>Crawl a docs site</h1>
<p>The crawler is one of three ways in…</p>
<h2 id="include-and-exclude-paths">Include and exclude paths</h2>
<p>Adjust which paths are crawled…</p>
</main>Wrap your main content region in data-suo-content to tell the crawler exactly where the real content starts, instead of guessing from the page's layout.
Per-URL states
While a crawl runs, each URL is one of:
- Queued — not yet fetched.
- Indexed — fetched and its records written.
- Skipped (with a reason) — matched an exclude rule, blocked by
robots.txt, or returned a non-indexable response (a redirect chain, anoindextag). - Failed (with a reason) — a network error, a timeout, or a non-2xx status.
The dashboard's page counter is always the live count of pages actually indexed so far, never a percentage estimate.
Re-crawling
suo re-crawls every verified site once a day, at 03:17 UTC. To pick up a deploy sooner, choose Start crawl on the Sources page, or trigger a crawl from CI after a deploy with suo crawl. Only one crawl of an index runs at a time: starting another while one is running answers crawl_running. A re-crawl only rewrites records whose content actually changed; unchanged pages keep their existing version and are not re-ranked.
Crawler identity and limits
The crawler only reads domains you have verified, identifies itself with its own user agent, obeys robots.txt and crawl-delay, and never sends credentials. See Security for the full crawler policy.