What does it do, and who is it for?

The tool does two jobs. It turns form fields into a valid robots.txt file with user-agent groups, ordered Allow and Disallow rules and Sitemap lines. It also evaluates a generated or pasted file and tells you whether a given path is allowed or blocked for a given bot, showing the group and the line that decided it. It is built for WordPress and WooCommerce developers, publishers setting an AI crawler policy, and SEOs reviewing crawl rules before a migration.

How to use it

  1. Pick a preset and press “Apply preset”. It replaces the rule groups; AI bot choices and sitemaps stay.
  2. In each group, enter comma-separated bot names, or * for every bot. For each rule choose Allow or Disallow and a path starting with /.
  3. In the AI bots section choose “No separate rule”, “Block” or “Allow” for each bot.
  4. Add sitemap URLs one per line and press “Generate robots.txt”. Read the warnings, then copy the output or download it as robots.txt.
  5. In the path tester, enter a bot name and a path such as /wp-admin/ or a full URL, then press “Test path”, or paste your live file to test it.

Group and path matching

The tester follows Google's robots.txt specification and labels its answer a “Google-compatible Allow/Disallow evaluation”. Groups whose user-agent matches the bot's product token, case-insensitively, are selected and merged. The * groups apply only when no specific group exists, so a bot with its own group ignores the * rules.

  • Rules match from the start of the path; * matches any sequence of characters and a trailing $ marks the end of the URL.
  • Among matching rules the longest one wins. If an Allow and a Disallow of equal length both match, Allow wins.
  • Paths are case-sensitive. Non-ASCII characters are compared as UTF-8 percent-encoding; %c3%96 equals %C3%96.
  • The query string is part of the path; a rule ending in $ does not cover a URL that continues with ?.
  • An empty Disallow: blocks nothing. Crawl-delay and unknown directives are listed by line but never affect the decision.

WordPress and WooCommerce presets

The WordPress preset blocks only the /wp-admin/ directory and allows /wp-admin/admin-ajax.php, which themes and plugins call from the front end. It leaves /wp-includes/ and /wp-content/ open because Google needs your CSS and JavaScript to render pages.

The WooCommerce preset adds the cart, checkout and account pages plus URLs carrying the add-to-cart parameter. Those page addresses depend on the install: a Turkish store might use /sepet/, /odeme/ or /hesabim/. Edit the fields to match the pages set in your WooCommerce settings.

Bot choices: search bots versus training bots

AI crawlers are not merged into one “block AI” switch because they do different jobs. OpenAI's GPTBot crawls for model training, while OAI-SearchBot powers ChatGPT search results. On the Anthropic side, ClaudeBot is for training and Claude-SearchBot for search quality. Google-Extended is not a separate crawler but a control token for Gemini training and grounding; it does not change your visibility in Google Search. PerplexityBot feeds Perplexity's search index; CCBot builds the Common Crawl archive used by many training datasets.

A bot with its own group ignores the * group, so choosing “Allow” does not carry your WordPress rules over to that bot. Names and purposes change; check the OpenAI and Anthropic documentation regularly.

Test examples

The table shows Googlebot results for the file below. Tested as GPTBot, the separate GPTBot group applies and Disallow: / blocks every path.

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /*.pdf$
Disallow: /Özel/

User-agent: GPTBot
Disallow: /
Sample results for Googlebot
PathResultWhy
/wp-admin/BlockedDisallow: /wp-admin/ matches
/wp-admin/admin-ajax.phpAllowedThe longer Allow rule wins
/x.pdfBlocked/*.pdf$ matches
/x.pdf?download=1Allowed$ does not cover a URL that continues with a query
/Özel/BlockedExact match
/özel/AllowedThe lower-case path is a different path

Crawling versus indexing, and common mistakes

robots.txt controls crawling, not indexing. A blocked URL can still appear in results, without a description, if other sites link to it. To remove a page from results, keep it crawlable and serve a noindex meta tag or an X-Robots-Tag header. robots.txt is not access control either: the file is public and badly behaved bots are free to ignore it. Google's robots.txt introduction explains the difference in detail.

  • A staging Disallow: / that reaches production closes the whole site; the tool warns strongly about it.
  • The tester is local and never downloads your live file; confirm with the robots.txt report in Search Console.
  • The file must live at the root of the host, at https://example.com/robots.txt; a file in a subdirectory is ignored.

Frequently asked questions

Where do I upload robots.txt?

Upload it to the root of your site so it opens at https://yourdomain.com/robots.txt. WordPress serves a virtual file until a real one exists. If an SEO plugin manages it, use the plugin's editor.

Does blocking GPTBot remove me from ChatGPT search?

No. ChatGPT search relies on OAI-SearchBot. Blocking only GPTBot leaves the search bot crawling under your other rules, which is why they are separate choices.

What is the difference between Disallow and noindex?

Disallow stops crawling; noindex keeps a crawled page out of the index. Google cannot see noindex on a page it may not crawl, so add noindex rather than blocking a page you want removed.

Is the test result identical to Google's behaviour?

It applies Google's published matching rules: group selection, longest match, * and $. It is still a local calculation that cannot see server responses, caching or redirects, and other bots may behave differently.

What does Crawl-delay do?

Google ignores Crawl-delay and sets its own crawl rate; some other bots may honour it. The tester reports the line but leaves it out of the decision.

Published: · Updated: