Technical readiness resource

robots.txt vs noindex: which one should you change?

Use crawl rules, indexing instructions, and authentication for their different jobs. Avoid removing the wrong protection while fixing search visibility.

robots.txt manages crawl access. noindex asks supporting search engines to exclude a page from results. Authentication protects private content. Choose the control that matches your goal.

If a public page should appear in search

Check for accidental noindex in the HTML robots meta tag and in the X-Robots-Tag response header. Also check whether the crawler can retrieve the page. A copied staging configuration can survive deployment, so inspect the production response directly instead of relying on local settings.

If a public page should stay out of results

Use a suitable indexing instruction and allow the search engine to read it. Blocking a URL in robots.txt can prevent it from seeing a noindex instruction on that page. A robots block alone is not a reliable removal mechanism.

<!-- For a page intentionally excluded from search results -->
<meta name="robots" content="noindex">

If the content is private

Require authentication and enforce authorization on the server. Neither a robots entry nor a noindex tag should be treated as a password. Test by opening the direct page URL in a fresh browser session. Admin dashboards, billing records, and saved user reports need actual access controls.

Use the report as a prompt to inspect intent

SiteReadyFor.SI flags a discovered noindex instruction and reports supported AI crawler rules. The right fix depends on whether the page is meant to be public. The report should help you notice an unexpected setting; it should not persuade you to expose deliberately private content.

References

Check your own website next.

Run an on-demand audit across crawler access, content retrieval, JSON-LD, sitemaps, and llms.txt.

Start a scan