AI visibility guide
Find the Firewall Rule Blocking an AI Search Crawler
Trace a refused AI crawler request from the edge to your origin, find the Cloudflare or host rule that matched, add a narrow exception, and confirm with real crawler traffic.
On this page
When an AI search crawler cannot fetch your WordPress pages even though robots.txt allows it, the refusal usually comes from a layer that actually answers requests: a CDN bot setting, a firewall rule, a rate limit, a host’s security layer, or a WordPress security plugin. Find which layer refused the request by comparing your edge logs with your origin logs, identify the exact rule from its event record, change only that rule, and confirm with the crawler’s real requests rather than a request you faked.
Separate policy from enforcement
Robots.txt and a firewall do different jobs. Robots.txt publishes your preferences; a firewall decides which requests are answered. The robots.txt standard, RFC 9309, checked September 27, 2026, says its rules “are not a form of access authorization.”
If you use WP Visibility, its AI crawler policy writes robots.txt rules only. Its documentation says it “does not block crawlers at the server” and points to your host’s or CDN’s bot rules for that; see AI crawler policy and llms.txt. A correct policy there cannot override a block at the edge.
Access matters for Google too, though Google’s AI features use ordinary Googlebot rather than a separate AI crawler. Google’s AI features documentation, checked September 27, 2026, says a supporting link must come from a page that is “indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements.”
Before tracing a block, confirm robots.txt is not the cause: the AI crawler policy guide covers reading the served file.
Separate edge failure from origin failure
Start with evidence of a real failure: a verified crawler request with a 403, 429, 503, or challenge response in your logs, or an operator’s tool or support reply telling you it could not fetch a page. Verify the request came from the operator’s published IP ranges first, as described in verifying AI bot visits. An unverified request with a crawler’s name tells you nothing about the real crawler.
Then look for the same request in two places:
| Found in edge security events | Found in origin access log | Where the block is |
|---|---|---|
| Yes, with a block, challenge, or rate limit action | No | The CDN or edge firewall |
| No | Yes, with a 403, 429, or 5xx status |
The host, the web server, or WordPress |
| No | No | Upstream of your logs: DNS, a host network layer, or the crawler never sent it |
| No | Yes, with 200 |
Not blocked at the time of that request |
Edge security events are not a full traffic log. Cloudflare’s security events documentation, checked September 27, 2026, says they cover requests its security products acted on or flagged, notes that its sampled logs may not list every event, and points to Security Analytics for all incoming traffic. Before concluding a request never reached the edge, look for it there. Cloudflare’s Security Analytics documentation, checked September 27, 2026, says it also shows whether traffic was served from Cloudflare’s network or from your origin server, which can explain a request missing from your origin log.
For the origin side, check in this order: your host’s control panel (many hosts run their own firewall or bot rules), web server rules such as .htaccess or server configuration, and WordPress security plugins with bot blocking or rate limiting. Your host’s support can tell you whether a platform-level rule exists that you cannot see.
A curl request with a crawler’s user agent is useful for one narrow question: whether a rule matches on the user agent string. It cannot establish whether the real crawler is allowed, because your request comes from your IP address and will not be treated as a verified bot.
Identify the matching Cloudflare rule
On Cloudflare, the event record usually names the product and rule responsible.
- Open Security Events. Cloudflare’s security events documentation, checked September 27, 2026, places it under the Analytics page, Events tab. Filter by the crawler’s user agent, the path, or the client IP address. Retention depends on plan: 24 hours on Free and Pro, 3 days on Business, 30 days on Enterprise, and the Free plan shows sampled logs only. Act quickly after a failure.
- Read the event. Note the action, the service (the product that acted), the rule, and the client IP. Check the IP against the operator’s list.
- Match the service to its setting:
| Service or setting | What to check | Documented behavior |
|---|---|---|
| AI Crawl Control | AI Crawl Control → Security → Crawlers, the crawler’s Actions column | Cloudflare’s AI Crawl Control documentation says blocking a crawler “creates or updates a WAF custom rule on your zone to enforce that block” |
| AI bot traffic policies | Search, Agent, and Training policies | Cloudflare’s AI traffic options changelog describes each as allow, block, or block only on pages with ads, with new defaults from September 15, 2026 for new domains |
| Block AI bots (legacy) | Security Settings → Block AI bots | Cloudflare’s Block AI bots page says it blocks bots crawling for AI training, marks it as deprecating on September 15, 2026, and says that from that date mixed-purpose crawlers that combine Search and Training are also blocked by it |
| Bot Fight Mode | Security Settings, bot traffic | Cloudflare’s Bot Fight Mode page says it issues challenges, is labeled Bot Fight Mode in the Service field, and that you “cannot bypass or skip Bot Fight Mode using WAF custom rules or Page Rules” |
| WAF custom rules | The Security rules page, custom rules | Rules you or a colleague wrote, or the rule AI Crawl Control created for a block, often matching user agent strings or countries |
| Rate limiting | Rate limiting rules | A crawler fetching many pages quickly can trip a limit meant for abuse |
Pages cited above were checked September 27, 2026. Dashboard labels change; use the event’s service field as the reliable pointer.
Two cases catch people out. A crawler that operates for both search and training can be affected by a training block; Cloudflare’s changelog says multi-purpose crawlers combining search and training are affected by defaults that block training. And a user agent rule written for one crawler can match another whose name contains the same text.
Retest a bounded exception
Change the one rule responsible, as narrowly as possible:
- AI Crawl Control: set that crawler to Allow on the Crawlers tab. Cloudflare’s documentation notes you can still choose to enforce robots.txt while allowing access.
- AI bot policies: if the crawler is classified as Search and your Search policy blocks it, decide whether that policy reflects your goals. If a mixed-purpose crawler is caught by a Training block, decide which matters more for that crawler.
- Your own custom rule: exclude the specific crawler from it, preferably by Cloudflare’s verified bot status where your plan exposes it, rather than by user agent text, which anyone can copy.
- Bot Fight Mode: because it cannot be skipped with custom rules, the choice is whether to keep it on for the zone, weighed against why it was turned on. For exceptions, Cloudflare’s Bot Fight Mode page points to Super Bot Fight Mode instead, which supports skip rules.
- Rate limiting: raise the threshold for the affected paths or exclude verified bots, rather than removing the limit.
- Origin rules: allow the operator’s published IP ranges in the host or plugin rule, and plan to refresh them; Perplexity’s crawler documentation, checked September 27, 2026, for example, recommends automated processes that periodically fetch its latest IP ranges from its endpoints.
Avoid broad fixes: turning off the firewall, allowing every request with a given user agent, or allowlisting a large IP block you have not matched to an operator’s published list.
Then confirm with real traffic:
- Record the time of the change.
- Wait for the crawler’s next requests. You usually cannot make a search crawler visit on demand, so allow days rather than minutes.
- Find its requests in edge and origin logs, verify the IPs, and confirm
200responses on the pages that previously failed. - Keep a short incident note: the failing request, the matching service and rule, the change, and the confirming request.
Checklist
- robots.txt ruled out, and the failing request verified against the operator’s IP ranges.
- The same request compared in edge events and origin logs.
- The responsible service and rule named from the event record.
- One narrow change made, not a broad allowlist.
- No exception based on user agent text alone.
- Real crawler requests confirmed with
200after the change, and the incident recorded.
