Robots.txt tester for search and AI crawlers
Paste a robots.txt and a URL path. The table shows, for each of the 14 search, training and user-triggered crawlers named in The SEO Playbook for 2027, whether it may fetch that path and the rule that decided. It runs entirely in your browser: nothing you paste leaves this page, and the page fetches nothing.
For /pricing: 10 of 14 allowed.
| Crawler | Used for | Result | Decided by |
|---|---|---|---|
| GooglebotGoogle · Search | Google Search, AI Overviews and AI Mode; also the Gemini app's grounding in Google Search.Google's common crawlers, updated 2026-07-14 | Allowed | No rule in the group naming Googlebot matches this path, so it is allowed |
| BingbotMicrosoft · Search | Bing, whose index Microsoft Copilot answers from.Bing Webmaster Guidelines, undated; revision reported 2026-02-26 | Allowed | No rule in the group naming Bingbot matches this path, so it is allowed |
| OAI-SearchBotOpenAI · Search | ChatGPT search. OpenAI says sites that opt out are not shown in ChatGPT search answers, though they can still appear as navigational links.Overview of OpenAI crawlers, undated | Allowed | No rule in the group naming OAI-SearchBot matches this path, so it is allowed |
| PerplexityBotPerplexity · Search | Perplexity's search answers, from its own crawl.Perplexity crawlers, modified 2026-01-29 | Allowed | No rule in the group naming PerplexityBot matches this path, so it is allowed |
| Claude-SearchBotAnthropic · Search | Claude's search results. Anthropic says blocking it may reduce a site's visibility there.Anthropic: crawling and how to block it, updated 2026-04-07 | Allowed | No rule in the group naming Claude-SearchBot matches this path, so it is allowed |
| ApplebotApple · Search | Siri, Spotlight and Safari, from Apple's own index.About Applebot, 2026-09-04 | Allowed | No rule in the group naming Applebot matches this path, so it is allowed |
| DuckAssistBotDuckDuckGo · Search | DuckDuckGo Assist answers (its links come "largely" from Bing).DuckAssistBot, undated | Allowed | No rule in the group naming DuckAssistBot matches this path, so it is allowed |
| GPTBotOpenAI · Training | Training OpenAI's models. OpenAI says it does not affect search.Overview of OpenAI crawlers, undated | Blocked | Disallow: /line 18, in the group naming GPTBot |
| ClaudeBotAnthropic · Training | Training Anthropic's models.Anthropic: crawling and how to block it, updated 2026-04-07 | Blocked | Disallow: /line 18, in the group naming ClaudeBot |
| Google-ExtendedGoogle · Training | Whether content trains and grounds Gemini models. Google says it does not affect Google Search or AI Overviews. A token, not a crawler: Googlebot does the fetching.Google's common crawlers, updated 2026-07-14 | Blocked | Disallow: /line 18, in the group naming Google-Extended |
| Applebot-ExtendedApple · Training | Whether Apple may use content for training its models. A token, not a crawler: Applebot does the fetching.About Applebot, 2026-09-04 | Blocked | Disallow: /line 18, in the group naming Applebot-Extended |
| ChatGPT-UserOpenAI · User-triggered | Fetches a page because a ChatGPT user asked for it.Overview of OpenAI crawlers, undated | Allowed | No rule in the User-agent: * group matches this path, so it is allowed |
| Perplexity-UserPerplexity · User-triggered | Fetches a page because a Perplexity user asked for it.Perplexity crawlers, modified 2026-01-29 | Allowed | No rule in the User-agent: * group matches this path, so it is allowed |
| Claude-UserAnthropic · User-triggered | Fetches a page because a Claude user asked for it.Anthropic: crawling and how to block it, updated 2026-04-07 | Allowed | No rule in the User-agent: * group matches this path, so it is allowed |
How the tester decides
It follows the robots.txt standard, RFC 9309 (September 2022), the way Google applies it:
- A crawler obeys only the groups that name it. Googlebot reads a
User-agent: Googlebotgroup and ignoresUser-agent: *entirely; only a crawler that no group names falls back to*. Names match without regard to case, and several groups naming the same crawler are merged. - The longest matching rule wins, so
Allow: /a/bbeatsDisallow: /afor/a/b/c. When an Allow and a Disallow match with equal length, Allow wins. *matches any characters, and a final$marks the end of the URL.Disallow: /*.pdf$blocks/a.pdfbut not/a.pdf?v=2. Rules match the path and query from the start.- No matching rule means allowed, and
/robots.txtitself is always allowed.
What a robots.txt test can't show
- Blocks in front of the file. A CDN or firewall can turn a crawler away before it ever reads robots.txt. Since 15 September 2026, Cloudflare's "Block" setting for AI crawlers also stops Googlebot, Bingbot and Applebot: see whether Cloudflare is blocking Googlebot. Cloudflare's managed robots.txt can also add lines to yours, so test the file as it is served, fetched from outside your network.
- Tokens are not crawlers. Google-Extended and Applebot-Extended control how content may be used; Googlebot and Applebot do the fetching. That is why blocking them leaves Search, AI Overviews and Apple's search features alone.
- User-triggered fetchers. ChatGPT-User, Claude-User and Perplexity-User fetch a page because a person asked. In the example they have no group of their own, so any of them that read robots.txt follow the
*group (Chapter 9).
Search and training are two separate decisions. Blocking a search crawler such as OAI-SearchBot or Claude-SearchBot can take you out of that engine's answers; blocking GPTBot, ClaudeBot, Google-Extended or Applebot-Extended is a business decision about model training, and the companies say it does not affect search (Chapter 9; sources in the table).
From the book
The example above is the worked robots.txt in Chapter 9, which is free to read: how to optimize your website for AI search. The full book, The SEO Playbook for 2027, covers robots.txt, CDN settings and crawl control in Chapters 2 and 8, with numbered steps that each end in a check.