Technical readiness resource

How to check robots.txt rules for AI crawlers

Find unintended crawler blocks, inspect rules for a specific page, and distinguish crawl permission from indexing and private access.

robots.txt communicates which paths a crawler may request. Check the rules for the specific bot and page you care about. An allow rule does not establish that the page was visited, indexed, or cited.

Check the exact address

Open /robots.txt at the root of your domain. If your application returns its homepage for every unknown path, the file may appear to load while actually containing HTML. Inspect the response body. Then choose one public URL, such as a service page, to evaluate against the rules.

Review broad rules before changing specific ones

A group for User-agent: * may apply broadly, while a more specific group targets a named crawler. Read the intended policy before removing a block. Different crawlers serve different purposes, so permission should reflect your site’s choices. Keep private data behind authentication regardless of robot rules.

# Illustrative only: permits GPTBot to crawl this site.
# Use only if that matches your intended policy.
User-agent: GPTBot
Allow: /

Sitemap: https://your-site.example/sitemap.xml

Check more than permission

If a rule permits access but the page still fails, inspect the HTTP response and firewall behavior. A login page, challenge screen, or unavailable server can prevent retrieval. Keep a record of the URL, response status, and rule you changed so a rescan has a clear comparison.

Use the free checker for a first pass

SiteReadyFor.SI’s robots tool evaluates its supported AI user agents against your supplied page URL. An audit adds content, schema, and sitemap findings. These checks help troubleshoot technical access; they do not certify that a crawler visited your website.

References

Check your own website next.

Run an on-demand audit across crawler access, content retrieval, JSON-LD, sitemaps, and llms.txt.

Start a scan