What does it check?
The checker reads your robots.txt and shows whether known AI crawlers may fetch the path you pick. You paste the text; the tool requests no address. robots.txt is required, llms.txt and page HTML are optional.
- Bot table: a verdict per bot, the matched group and the deciding rule with its line number.
- Summary: training bots blocked and AI search bots allowed; user-triggered fetchers and control tokens are counted separately.
- llms.txt format, unsafe links; in HTML: JSON-LD, meta robots, canonical, title, description, lang.
The result is neutral: blocking training bots while allowing search bots is a valid choice, as is allowing everything. The site URL field only builds links to robots.txt and llms.txt for a new tab.
Training, search and user-triggered bots
One company can run several bots with different jobs. Operator documentation, rechecked on 2026-10-11, separates three purposes.
| Purpose | Tokens | Operator documentation |
|---|---|---|
| Training | GPTBot, ClaudeBot, CCBot, Amazonbot, MistralAI-Training | Training-related crawling; robots.txt is the documented opt-out for GPTBot, CCBot and MistralAI-Training. Amazonbot may train Amazon AI. |
| Search | OAI-SearchBot, Claude-SearchBot, PerplexityBot, Amzn-SearchBot, DuckAssistBot, MistralAI-Index | Search and answer features; PerplexityBot, DuckAssistBot and MistralAI-Index are described as not training. |
| User-triggered | ChatGPT-User, Claude-User, Perplexity-User, Amzn-User, MistralAI-User | Fetches a page when a person asks. OpenAI says robots.txt may not apply; Perplexity generally ignores it; Anthropic says Claude-User respects it. |
Google-CloudVertexBot is a site-owner-requested Vertex AI crawl outside the totals. Meta's three tokens and Bytespider show “unknown” and join no total; Googlebot is a reference row.
How do robots.txt rules match?
Matching follows RFC 9309. Lines group by User-agent: consecutive lines share a group and groups for one token merge. A bot uses the groups matching its token exactly (case-insensitive); only if none exist does the * group apply.
- Paths are case-sensitive;
*is a wildcard,$anchors the end. - The longest matching pattern wins; on a tie
Allowwins. - An empty
Disallow:allows everything. - Percent-encoding is normalised before comparison.
User-agent: *
Disallow: /private/
Allow: /private/public/
User-agent: GPTBot
Disallow: /GPTBot uses only its own group and is blocked at /; OAI-SearchBot falls back to * and is allowed. Other Disallow rules alone do not prove other paths are blocked, since an Allow may override them, so they are reported as information. Operators describe fallbacks to other bots' rules for Google-CloudVertexBot and Amzn-SearchBot; the generic matcher does not reproduce them, so those rows carry a caveat.
Why are Google-Extended and Applebot-Extended “control tokens”?
These names are robots.txt tokens, not separate crawlers. Google says Google-Extended manages use for Gemini training and grounding and does not affect Search. Apple says Applebot-Extended does not crawl pages and controls training use.
So these rows read “usage permitted” or “usage excluded”, not “allowed” or “blocked”: they show a preference about use, not crawl access. Sources: Google and Apple.
What llms.txt is and is not
llms.txt is a proposal by Jeremy Howard (3 September 2024), not a standard. llmstxt.org showed “version 2”, last modified 10 August 2026, when accessed on 11 October 2026. Only an H1 at the top is mandatory; the blockquote summary and H2 link sections are suggestions, not format requirements.
No verified source says llms.txt improves ranking or AI visibility, and the tool claims none. It reports a missing H1, unsafe link schemes (javascript:, data:, //) and size. Relative links are listed as unresolved context, never clickable. To write the file, use the llms.txt generator.
Structured data and meta robots
With HTML pasted, the tool counts JSON-LD blocks, parses each and lists @type values; invalid JSON is an error. Use the JSON-LD validator for detail and the schema generator for markup. Structured data is not promised to improve AI visibility; the tool only shows whether it exists.
It looks for noindex, nosnippet, max-snippet and noarchive in meta name="robots" and googlebot, plus canonical, title, meta description and lang. A robots.txt block is not noindex: Google says blocking crawling can prevent content retrieval but does not guarantee URL removal from Search (Google documentation). If Googlebot is blocked, it may never see a noindex.
Limits
- Paste mode: nothing is downloaded, so the live file may differ from what you pasted.
- robots.txt is voluntary; a bot that ignores it is not stopped.
- CDN, WAF or server-level bot rules are invisible here.
- A rule match is not proof a request was blocked; operators may enforce differently for user-triggered fetchers.
- Limits: robots.txt 500 KiB and 20,000 lines (RFC 9309 asks for at least 500 KiB); llms.txt 500 KiB and HTML 2 MiB are tool limits.
- HTML is never executed, displayed or loaded, and never enters a share link.
Frequently asked questions
Do I have to block AI bots?
No. Blocking training but allowing search, allowing all or blocking all are valid; the tool does not judge. Documentation: OpenAI, Anthropic, Google, Perplexity.
Does robots.txt stop a bot for certain?
No. It is voluntary and operators differ: OpenAI says rules may not apply to ChatGPT-User; Perplexity says Perplexity-User generally ignores them. Reliable control needs server-side methods.
What does “unknown” mean?
Meta's three tokens and Bytespider lack verified official documentation (Meta's page was unavailable in this audit; none was found for Bytespider), so they get no verdict and join no total.
Will llms.txt improve my AI visibility?
No verified source says so and the tool does not promise it. llms.txt is a proposal; no search engine or AI company must read it.
Is my pasted content sent anywhere?
No. Everything runs in your browser; nothing is downloaded and pasted HTML loads no images or frames. Share links carry no input by default; robots.txt and llms.txt are added only if you tick the box, never HTML.