Skip to content

ToolsAI crawler access check

AI crawler access check

See whether named AI crawlers are allowed to read your site, and whether they can actually reach the homepage.

How this tool works

AI assistants only cite pages they can fetch. A robots.txt rule or a WAF that returns 403 is enough to keep you out of the answer.

This check reads your robots.txt for every named AI agent, then requests the homepage as that agent. It also looks for llms.txt and flags pages whose main copy looks as if it only exists after JavaScript.

The result is a pass or fail per agent, with the exact change for each failure. It is a snapshot of today, not a trend.

Questions

What does the AI crawler access check test?

It parses robots.txt for each named AI agent, including wildcard rules, then requests the homepage with that agent's user-agent. It also checks for llms.txt and looks at whether the HTML we received has readable content or looks like a JavaScript shell.

Why can robots.txt allow a bot that still gets a 403?

robots.txt is advice. A CDN or WAF can still block the request. That is the common failure this check is for: the file says Allow, the homepage returns 403, and the assistant never reads the page.

Does this run a real browser?

No. The JavaScript note is inferred from the HTML we received: empty roots, script-heavy pages, or copy that only appears in a noscript tag. It is not a headless render. The live fetch per agent is a real HTTP request.

Is the full result free?

Yes. Every agent, every status code, and every suggested fix is on the page. There is no email gate on this tool.