Practical guide

Website crawl sources: pages, sitemaps, links, and exclusions

A crawler may discover pages through internal links, sitemaps, or configured lists. Comparing discovery paths exposes orphan pages, stale URLs, blocked sections, and incomplete navigation. Use current primary documentation, a representative project, explicit definitions, and a recoverable operating plan before scaling the workflow.

Last materially reviewed 2026-08-28

Quick answerA crawler may discover pages through internal links, sitemaps, or configured lists. Comparing discovery paths exposes orphan pages, stale URLs, blocked sections, and incomplete navigation.
Direct answer

Crawl sources: the working definition

A crawler may discover pages through internal links, sitemaps, or configured lists. Comparing discovery paths exposes orphan pages, stale URLs, blocked sections, and incomplete navigation.

Start with the business or client decision that crawl sources must improve, then identify the exact project, market, page, metric, owner, and review date. That keeps crawl sources attached to an operating job instead of a dashboard habit. A useful conclusion states both the action and the evidence that could reverse it.

  • Define what success means for crawl sources.
Working method

Translate the metric into a decision

Build evidence in layers. Product documentation defines supported controls; Google data establishes search observations; analytics establishes on-site outcomes where configured; and the team’s change log explains what actually shipped.

Each source has one job. A visibility score cannot prove revenue, and a crawl warning cannot set business priority without affected pages and consequences.

  • Name the source and scope.
  • Assign the workflow owner.
  • Record the fact that would reverse the decision.
Decision framework

Avoid the common interpretation error

Use a quality gate with five rows: source integrity, configuration fit, decision clarity, action ownership, and verification. Mark any row that relies on an unexplained metric, an inaccessible account, a stale export, or one person’s memory.

A workflow stays active only when every material row has evidence and an owner.

  • Compare the same operating job.
  • Keep cost and review effort in the model.
  • Preserve a recoverable fallback.
Final check

Create the operating record

Commission crawl sources like a production process. Capture screenshots, export a baseline, test recipient access, trigger one expected alert, simulate one failure, verify the fallback, and record the support route.

The final deliverable is not a configured screen. It is a repeatable decision with an owner and recovery path.

  • Save settings and definitions.
  • Test one complete cycle.
  • Schedule the next review.
Continue when useful

Next: Audit setup

Domain variant, scope, robots behavior, sitemap, crawl sources, limits, authentication, user agent, exclusions, and schedule determine what the audit can see. Capture the exact crawl timestamp and deployment state so a later difference has a plausible reference point. Use current primary documentation, a representative project, explicit definitions, and a recoverable operating plan before scaling the workflow.

Open Audit setup →

Sources used for this page

These records support the facts and comparisons above. Merchant-controlled records are labelled so you can separate product claims from independent evidence.

  1. How to use Website Audit — DOCUMENTATION · checked 2026-08-28
  2. Google Search Essentials and SEO Starter Guide — DOCUMENTATION · checked 2026-08-28