AI visibility guide
Allow AI Search Crawlers Without Opening Every Training Bot
Map each AI provider's documented crawlers to search, user request, and training roles, write a robots.txt policy that allows search but not training, and test what is served.
On this page
OpenAI, Anthropic, and Perplexity each publish separate crawler names for search, for fetches a user asks for, and (except Perplexity) for model training, and Google has a separate token for Gemini training. To stay citable in AI search while opting out of training, disallow the training crawlers by name in robots.txt, leave the search crawlers allowed, and then check the robots.txt your site actually serves, including the wildcard group and anything your CDN adds. Robots.txt records a preference; whether each operator honors it is their documented policy, not something the file enforces.
Define the site’s publication goals
Decide what you want before you write rules. Three questions cover most sites:
- Do you want to appear, with links, in AI search answers? For most businesses and publishers selling their own products or services, yes.
- Do you want your content used to train future models? A separate decision. Some sites are indifferent; some license their content and want to keep it out.
- Do you want an assistant to be able to open a page when a user asks it to? Blocking this asks assistants not to make those live fetches; Anthropic says (checked September 27, 2026) blocking Claude-User “may reduce your site’s visibility for user-directed web search.”
Write the answers down with a date. They are the policy; the robots.txt lines only implement it.
A training opt-out looks forward. Anthropic (checked September 27, 2026), for example, says restricting ClaudeBot signals that a site’s “future materials” should be excluded from its training datasets; a robots.txt rule does not remove content already collected. The general explanation of AI visibility, and what access can and cannot establish, is in what AI visibility means.
Map documented bot roles
The table below summarizes each provider’s own documentation, checked September 27, 2026.
| Provider | Search or retrieval | User-requested fetch | Training |
|---|---|---|---|
| OpenAI | OAI-SearchBot | ChatGPT-User | GPTBot |
| Anthropic | Claude-SearchBot | Claude-User | ClaudeBot |
| Perplexity | PerplexityBot | Perplexity-User | None listed |
| Googlebot (Search, including AI features) | Not covered here | Google-Extended token |
What each provider says:
- OpenAI. Its crawler documentation describes OAI-SearchBot as “used to surface websites in search results in ChatGPT’s search features” and GPTBot as used “to make our generative AI foundation models more useful and safe.” It says “each setting is independent of the others.” For ChatGPT-User, it notes that “because these actions are initiated by a user, robots.txt rules may not apply.” It also says that for search results it “can take ~24 hours from a site’s robots.txt update” for its systems to adjust.
- Anthropic. Its crawler help article lists ClaudeBot as collecting web content that could contribute to training its models, Claude-SearchBot as improving “search result quality for users,” and Claude-User as accessing sites “when individuals ask questions to Claude,” and says what disabling each one does. Its robots.txt disallow example uses ClaudeBot; the same form works for the other two names.
- Perplexity. Its crawler documentation says PerplexityBot is “designed to surface and link websites in search results on Perplexity” and “is not used to crawl content for AI foundation models.” It says Perplexity-User “generally ignores robots.txt rules” because it acts on a user’s request.
- Google. Google’s list of common crawlers describes Google-Extended as a robots.txt token with no user agent of its own, controlling use of content for training Gemini models and for grounding in Gemini Apps and Vertex AI. It states that Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” Google’s AI Overviews and AI Mode are part of Search; its AI features guide, checked September 27, 2026, points to
nosnippet,data-nosnippet,max-snippet, andnoindexfor limiting what is shown, covered in limiting Google AI answer previews.
Note the Google-Extended trade-off: it covers grounding in Gemini as well as training, so disallowing it is not purely a training decision.
For the goal “search yes, training no,” the rules follow directly:
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /
Search crawlers get no group of their own and fall back to the wildcard group. That fallback is defined in the robots.txt standard, RFC 9309, checked September 27, 2026: “If no matching group exists, crawlers MUST obey the group with a user-agent line with the ‘*’ value, if present.” The same standard says: “These rules are not a form of access authorization.”
Set the policy in WordPress
If you use WP Visibility, you do not need to hand-edit the file. Under WP Visibility → Settings → AI Visibility, set Model training crawlers to Block and AI search crawlers to Allow. The plugin then appends a single group naming its training crawlers with Disallow: /. Its training list includes GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, meta-externalagent, and Bytespider; its AI search list includes OAI-SearchBot, Claude-SearchBot, PerplexityBot, ChatGPT-User, Claude-User, and DuckAssistBot. Per-bot overrides change one crawler without changing its group. The plugin has no separate user-fetch group, so use an override if your answer for ChatGPT-User or Claude-User differs from your search answer. The full field reference and the crawler lists are in AI crawler policy and llms.txt.
Two conditions from that doc override your settings: a physical robots.txt file in the site’s root folder is served instead of WordPress’s generated file, and while Discourage search engines from indexing this site under Settings → Reading is ticked, the plugin adds nothing to robots.txt, so none of these crawler rules are published. WordPress itself does not block crawlers in robots.txt on such a site; since version 5.3 it uses a noindex robots meta tag instead, per the do_robots() and wp_robots_no_robots() references, checked September 27, 2026.
Check actual robots responses
The settings screen shows what you intended. The file a crawler downloads shows what happens. Check it from outside:
curl -s https://example.com/robots.txt
curl -sI https://example.com/robots.txt | head -n 1
Use your own domain. Read the output with these questions:
| Check | What to look for | If it fails |
|---|---|---|
| Status | 200 |
Anything else needs investigating before you rely on the rules |
| Wildcard group | What User-agent: * disallows |
A Disallow: / here disallows every search crawler that has no group of its own |
| Training group | GPTBot, ClaudeBot, Google-Extended, and any others you chose, with Disallow: / |
Missing names are not opted out |
| Search crawlers | No group that disallows OAI-SearchBot, Claude-SearchBot, or PerplexityBot | A per-bot override or old manual rule may be blocking one |
| Extra content | Lines you did not write | A CDN or host may be adding rules |
CDN-added rules. Cloudflare’s managed robots.txt documentation, checked September 27, 2026, says that when the setting is on and the site already serves a robots.txt, “Cloudflare will prepend our managed robots.txt before your existing robots.txt, combining both into a single response.” Its managed section disallows known AI crawlers and adds a default content signal. If you see blocks you did not write, check that setting.
Bot controls beyond robots.txt. A correct file does not help if the edge blocks the request. Cloudflare’s changelog for AI traffic options, checked September 27, 2026, describes separate Search, Agent, and Training policies, each set to allow, block, or block only on pages with ads, with new defaults from September 15, 2026 for new domains that block Training and Agent on pages that display ads while Search remains allowed. Review those settings alongside your robots.txt. If a search crawler is being refused, the firewall rule guide traces it.
Separate stated compliance from observed behavior
The provider documentation above tells you what each operator says it does. To see what actually happens on your site, look at your server or CDN logs after the change:
- Allowed search crawlers should continue to fetch pages with
200responses. - Disallowed training crawlers should stop requesting pages other than robots.txt, allowing time for each operator to fetch the updated file. Google-Extended has no user agent of its own, so it never appears in logs.
- A user agent string can be copied by anyone. Before concluding that a provider ignored your file, confirm the requests came from its published IP ranges, as described in verifying AI bot visits.
Record the date of each policy change and what the logs show a week or two later. That record answers “did it work” better than the settings screen.
Checklist
- Written goals for search, user fetches, and training, with a date.
- Training crawlers disallowed by name; search crawlers not disallowed.
- The Google-Extended grounding trade-off considered.
- Served robots.txt checked, including the wildcard group.
- CDN-managed robots.txt and AI bot policies reviewed.
- Logs checked after the change, with claimed bots verified by IP.
