Skip to content
WP Visibility

AI visibility guide

Find the Firewall Rule Blocking an AI Search Crawler

Trace a refused AI crawler request from the edge to your origin, find the Cloudflare or host rule that matched, add a narrow exception, and confirm with real crawler traffic.

Published

On this page

When an AI search crawler cannot fetch your WordPress pages even though robots.txt allows it, the refusal usually comes from a layer that actually answers requests: a CDN bot setting, a firewall rule, a rate limit, a host’s security layer, or a WordPress security plugin. Find which layer refused the request by comparing your edge logs with your origin logs, identify the exact rule from its event record, change only that rule, and confirm with the crawler’s real requests rather than a request you faked.

Separate policy from enforcement

Robots.txt and a firewall do different jobs. Robots.txt publishes your preferences; a firewall decides which requests are answered. The robots.txt standard, RFC 9309, checked September 27, 2026, says its rules “are not a form of access authorization.”

If you use WP Visibility, its AI crawler policy writes robots.txt rules only. Its documentation says it “does not block crawlers at the server” and points to your host’s or CDN’s bot rules for that; see AI crawler policy and llms.txt. A correct policy there cannot override a block at the edge.

Access matters for Google too, though Google’s AI features use ordinary Googlebot rather than a separate AI crawler. Google’s AI features documentation, checked September 27, 2026, says a supporting link must come from a page that is “indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements.”

Before tracing a block, confirm robots.txt is not the cause: the AI crawler policy guide covers reading the served file.

Where a crawler can be refused. robots.txt: A published preference, not enforcement. CDN edge: Bot settings, firewall rules, rate limits. Origin: Host rules, server config, security plugins.
A request can pass robots.txt and still be refused at the edge or the origin. Compare edge events with origin logs to find the layer that acted.

Separate edge failure from origin failure

Start with evidence of a real failure: a verified crawler request with a 403, 429, 503, or challenge response in your logs, or an operator’s tool or support reply telling you it could not fetch a page. Verify the request came from the operator’s published IP ranges first, as described in verifying AI bot visits. An unverified request with a crawler’s name tells you nothing about the real crawler.

Then look for the same request in two places:

Found in edge security events Found in origin access log Where the block is
Yes, with a block, challenge, or rate limit action No The CDN or edge firewall
No Yes, with a 403, 429, or 5xx status The host, the web server, or WordPress
No No Upstream of your logs: DNS, a host network layer, or the crawler never sent it
No Yes, with 200 Not blocked at the time of that request

Edge security events are not a full traffic log. Cloudflare’s security events documentation, checked September 27, 2026, says they cover requests its security products acted on or flagged, notes that its sampled logs may not list every event, and points to Security Analytics for all incoming traffic. Before concluding a request never reached the edge, look for it there. Cloudflare’s Security Analytics documentation, checked September 27, 2026, says it also shows whether traffic was served from Cloudflare’s network or from your origin server, which can explain a request missing from your origin log.

For the origin side, check in this order: your host’s control panel (many hosts run their own firewall or bot rules), web server rules such as .htaccess or server configuration, and WordPress security plugins with bot blocking or rate limiting. Your host’s support can tell you whether a platform-level rule exists that you cannot see.

A curl request with a crawler’s user agent is useful for one narrow question: whether a rule matches on the user agent string. It cannot establish whether the real crawler is allowed, because your request comes from your IP address and will not be treated as a verified bot.

Identify the matching Cloudflare rule

On Cloudflare, the event record usually names the product and rule responsible.

  1. Open Security Events. Cloudflare’s security events documentation, checked September 27, 2026, places it under the Analytics page, Events tab. Filter by the crawler’s user agent, the path, or the client IP address. Retention depends on plan: 24 hours on Free and Pro, 3 days on Business, 30 days on Enterprise, and the Free plan shows sampled logs only. Act quickly after a failure.
  2. Read the event. Note the action, the service (the product that acted), the rule, and the client IP. Check the IP against the operator’s list.
  3. Match the service to its setting:
Service or setting What to check Documented behavior
AI Crawl Control AI Crawl Control → Security → Crawlers, the crawler’s Actions column Cloudflare’s AI Crawl Control documentation says blocking a crawler “creates or updates a WAF custom rule on your zone to enforce that block”
AI bot traffic policies Search, Agent, and Training policies Cloudflare’s AI traffic options changelog describes each as allow, block, or block only on pages with ads, with new defaults from September 15, 2026 for new domains
Block AI bots (legacy) Security Settings → Block AI bots Cloudflare’s Block AI bots page says it blocks bots crawling for AI training, marks it as deprecating on September 15, 2026, and says that from that date mixed-purpose crawlers that combine Search and Training are also blocked by it
Bot Fight Mode Security Settings, bot traffic Cloudflare’s Bot Fight Mode page says it issues challenges, is labeled Bot Fight Mode in the Service field, and that you “cannot bypass or skip Bot Fight Mode using WAF custom rules or Page Rules”
WAF custom rules The Security rules page, custom rules Rules you or a colleague wrote, or the rule AI Crawl Control created for a block, often matching user agent strings or countries
Rate limiting Rate limiting rules A crawler fetching many pages quickly can trip a limit meant for abuse

Pages cited above were checked September 27, 2026. Dashboard labels change; use the event’s service field as the reliable pointer.

Two cases catch people out. A crawler that operates for both search and training can be affected by a training block; Cloudflare’s changelog says multi-purpose crawlers combining search and training are affected by defaults that block training. And a user agent rule written for one crawler can match another whose name contains the same text.

Retest a bounded exception

Change the one rule responsible, as narrowly as possible:

  • AI Crawl Control: set that crawler to Allow on the Crawlers tab. Cloudflare’s documentation notes you can still choose to enforce robots.txt while allowing access.
  • AI bot policies: if the crawler is classified as Search and your Search policy blocks it, decide whether that policy reflects your goals. If a mixed-purpose crawler is caught by a Training block, decide which matters more for that crawler.
  • Your own custom rule: exclude the specific crawler from it, preferably by Cloudflare’s verified bot status where your plan exposes it, rather than by user agent text, which anyone can copy.
  • Bot Fight Mode: because it cannot be skipped with custom rules, the choice is whether to keep it on for the zone, weighed against why it was turned on. For exceptions, Cloudflare’s Bot Fight Mode page points to Super Bot Fight Mode instead, which supports skip rules.
  • Rate limiting: raise the threshold for the affected paths or exclude verified bots, rather than removing the limit.
  • Origin rules: allow the operator’s published IP ranges in the host or plugin rule, and plan to refresh them; Perplexity’s crawler documentation, checked September 27, 2026, for example, recommends automated processes that periodically fetch its latest IP ranges from its endpoints.

Avoid broad fixes: turning off the firewall, allowing every request with a given user agent, or allowlisting a large IP block you have not matched to an operator’s published list.

Then confirm with real traffic:

  1. Record the time of the change.
  2. Wait for the crawler’s next requests. You usually cannot make a search crawler visit on demand, so allow days rather than minutes.
  3. Find its requests in edge and origin logs, verify the IPs, and confirm 200 responses on the pages that previously failed.
  4. Keep a short incident note: the failing request, the matching service and rule, the change, and the confirming request.

Checklist

  • robots.txt ruled out, and the failing request verified against the operator’s IP ranges.
  • The same request compared in edge events and origin logs.
  • The responsible service and rule named from the event record.
  • One narrow change made, not a broad allowlist.
  • No exception based on user agent text alone.
  • Real crawler requests confirmed with 200 after the change, and the incident recorded.

Read next

WordPress SEO with your own assistant.

WP Visibility is $99 a year for unlimited sites, client sites included, with a 30-day refund. Use its SEO tools in WordPress or connect a supported assistant. Read how proposal review and permissions work.