AI visibility guide
Verify AI Bot Visits Before You Trust a Crawler Report
Check claimed AI crawler requests in your logs against each operator's published IP ranges or DNS method, and report verified, unverified, and unknown traffic separately.
On this page
A user agent string that says “GPTBot” or “ClaudeBot” is a claim, not proof. To verify an AI crawler visit, take the request’s IP address from your logs and check it against the method the operator publishes: IP range files for OpenAI, Anthropic, and Perplexity, and a reverse and forward DNS lookup (or IP range files) for Google. Then report three groups separately: verified, claimed but unverified, and unclaimed. Any crawler report that counts user agents alone is counting claims.
Know why a user agent is not enough
Any HTTP client can send any user agent. Cloudflare’s verified bots documentation, checked September 27, 2026, reflects this: to treat a bot as verified, Cloudflare requires honest self-identification through a cryptographic Web Bot Auth signature, a published IP list with a stable user agent, or reverse DNS. The user agent alone is not one of those methods.
This matters in both directions. A spoofed “GPTBot” can inflate a report of AI crawler activity. It can also make it look as if an operator ignored your robots.txt when the requests came from someone else.
If you use WP Visibility’s optional AI Traffic Lens module (off by default in 2.10.3), know what it measures. It counts requests whose user agent contains a known AI crawler name, plus visits referred from assistants, and it stores no IP addresses. Its counts are therefore claimed visits. It also sees only requests that reach WordPress; a page served from a full-page cache or a CDN never runs the plugin. The command that prints its report is in the WP-CLI reference.
Collect a sanitized request sample
Work from raw access logs: your host’s logs, your web server’s, or your CDN’s. Pick a fixed window, such as seven days.
- Find where the real client IP is. If your site sits behind a CDN or proxy, your origin’s logs may show the proxy’s addresses. Use the CDN’s logs, or the request header your CDN documents for the original client IP.
- Extract claimed AI crawler requests. For a standard combined log format:
grep -E "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|Googlebot" access.log > claimed.log
awk '{print $1}' claimed.log | sort -u > claimed-ips.txt
Adjust the file names and the field number to your log format.
- Keep the working files private. IP addresses in logs can identify people, especially in spoofed requests from ordinary connections. Do not paste raw logs into shared documents or tickets. For anything you share, replace IP addresses with the verification result, or truncate them.
- Add a deliberate spoof case. From your own connection, request a page while claiming to be a crawler:
curl -s -o /dev/null -A "Mozilla/5.0 (compatible; GPTBot/1.4; +https://openai.com/gptbot)" https://example.com/
Use your own domain. Your request should appear in the logs with a GPTBot user agent and your IP address. It must come out of the next step as unverified; if it does not, the check is not working.
Apply supported verification methods
Each operator documents its own method.
| Operator | Published method | Source, checked September 27, 2026 |
|---|---|---|
| OpenAI | Separate IP range files for OAI-SearchBot, GPTBot, ChatGPT-User, and OAI-AdsBot | OpenAI crawler documentation |
| Anthropic | One IP range file for its bots | Anthropic crawler help article |
| Perplexity | Separate IP range files for PerplexityBot and Perplexity-User | Perplexity crawler documentation |
| Reverse then forward DNS lookup, or IP range files | Verify requests from Google crawlers and fetchers |
Anthropic’s article says that if a crawler “has a source IP address on this list, it indicates that the crawler is coming from Anthropic.” Because the article publishes a single list rather than one per bot, an IP match confirms Anthropic but not which of its bots made the request. Perplexity’s documentation notes that WAF changes “may take some time to propagate” and recommends refreshing IP ranges from its endpoints periodically; the same applies to any list you download.
Check IP ranges with a script. This Python script reads IP addresses, one per line, and reports which published range contains each. It uses only the standard library.
import ipaddress, json, sys, urllib.request
SOURCES = {
"OpenAI OAI-SearchBot": "https://openai.com/searchbot.json",
"OpenAI GPTBot": "https://openai.com/gptbot.json",
"OpenAI ChatGPT-User": "https://openai.com/chatgpt-user.json",
"Anthropic (any of its bots)": "https://claude.com/crawling/bots.json",
"PerplexityBot": "https://www.perplexity.com/perplexitybot.json",
"Perplexity-User": "https://www.perplexity.com/perplexity-user.json",
}
def load(url):
req = urllib.request.Request(url, headers={"User-Agent": "range-check"})
with urllib.request.urlopen(req, timeout=20) as response:
data = json.load(response)
networks = []
for prefix in data.get("prefixes", []):
cidr = prefix.get("ipv4Prefix") or prefix.get("ipv6Prefix")
if cidr:
networks.append(ipaddress.ip_network(cidr))
return networks
ranges = {name: load(url) for name, url in SOURCES.items()}
for line in sys.stdin:
ip = line.strip()
if not ip:
continue
address = ipaddress.ip_address(ip)
owners = [name for name, nets in ranges.items() if any(address in net for net in nets)]
print(ip, "->", ", ".join(owners) or "no match")
Run it with python check_ranges.py < claimed-ips.txt. The file URLs are the ones each operator’s documentation lists; if a URL stops working, return to that documentation rather than guessing a replacement.
Check Google with DNS. Google’s page gives four steps: run a reverse DNS lookup on the IP with the host command, confirm the name ends in googlebot.com, google.com, or googleusercontent.com, run a forward lookup on that name, and confirm it returns the original IP.
host <ip-from-your-log>
host <name-returned-above>
Both lookups must agree. A reverse lookup alone can be set by whoever controls the IP’s DNS.
Report verified and uncertain traffic
Match each request’s claimed identity with the verification result:
| Claimed user agent | Verification result | Label |
|---|---|---|
| GPTBot | IP in OpenAI’s GPTBot range | Verified GPTBot |
| GPTBot | IP in no OpenAI range | Claimed, unverified |
| ClaudeBot | IP in Anthropic’s range | Verified Anthropic (bot not distinguished by IP) |
| Googlebot | Reverse and forward DNS agree on googlebot.com | Verified Google |
| Any AI crawler name | IP matches a different operator’s range | Inconsistent; investigate |
| No crawler name | Any | Unclaimed, outside this report |
Your spoof request should be in the second row.
A useful report states:
- The log source and time window.
- Verified requests per operator and bot, with status codes. A verified crawler receiving
403or429responses is an access problem; the firewall rule guide traces it. - Claimed but unverified requests, as a separate total. Do not add them to the operator’s figures.
- Anything you could not check, such as requests with no usable client IP.
- The date the IP lists were downloaded.
Do not conclude that an operator ignored robots.txt from unverified requests. Check the verified group against your rules, allowing for the adjustment periods operators mention, such as OpenAI’s note that, for search results, it can take about 24 hours after a robots.txt update for its systems to adjust (OpenAI crawler documentation, checked September 27, 2026). The policy side is in the AI crawler policy guide.
Checklist
- Logs contain the real client IP, not a proxy’s.
- A spoofed request was added and came out unverified.
- IP ranges downloaded fresh, with the date recorded.
- Google requests checked with reverse and forward DNS.
- Verified, unverified, and unclaimed traffic reported separately.
- No raw visitor IP addresses in anything shared.
