Technical readiness resource

Why are pages missing from your sitemap or site scan?

Compare your public URL inventory with your sitemap, check nested sitemap files, and verify that new content is discoverable after publishing.

A sitemap only helps with the URLs it actually lists. Missing entries, invalid XML, a stale generator, or a discovery limit can leave pages out of a scan. A sitemap is a discovery aid, not a promise of indexing.

Build an inventory from the source of your content

List the public pages from your CMS, product database, or route records. Compare that list with the sitemap rather than judging completeness from a successful HTTP response. For example, a store with twenty published product records should account for twenty intended product URLs or explain why some are excluded.

Inspect the actual XML

Open /sitemap.xml and confirm it contains XML rather than an HTML fallback. If it is a sitemap index, open its child files. Prefer the canonical public addresses you want discovered and remove deleted or redirected destinations.

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url><loc>https://your-site.example/</loc></url>
  <url><loc>https://your-site.example/services</loc></url>
</urlset>

Publish a new page through the whole workflow

Add one test article to your content source, publish it, and check its direct URL. Confirm the generator includes it in the sitemap and that a relevant public page links to it. If the content is generated from records, the sitemap should use the same published records so the lists do not drift apart.

Interpret SiteReadyFor.SI’s result correctly

SiteReadyFor.SI starts at /sitemap.xml, follows bounded same-origin sitemap indexes, and caps discovery at 500 pages. A valid sitemap score does not prove complete site coverage. A discovery-limit message may explain a partial scan; compare the found list with your inventory before deciding a page is missing from search.

References

Check your own website next.

Run an on-demand audit across crawler access, content retrieval, JSON-LD, sitemaps, and llms.txt.

Start a scan