Skip to content
WP Visibility

AI visibility guide

Verify AI Bot Visits Before You Trust a Crawler Report

Check claimed AI crawler requests in your logs against each operator's published IP ranges or DNS method, and report verified, unverified, and unknown traffic separately.

Published

On this page

A user agent string that says “GPTBot” or “ClaudeBot” is a claim, not proof. To verify an AI crawler visit, take the request’s IP address from your logs and check it against the method the operator publishes: IP range files for OpenAI, Anthropic, and Perplexity, and a reverse and forward DNS lookup (or IP range files) for Google. Then report three groups separately: verified, claimed but unverified, and unclaimed. Any crawler report that counts user agents alone is counting claims.

Know why a user agent is not enough

Any HTTP client can send any user agent. Cloudflare’s verified bots documentation, checked September 27, 2026, reflects this: to treat a bot as verified, Cloudflare requires honest self-identification through a cryptographic Web Bot Auth signature, a published IP list with a stable user agent, or reverse DNS. The user agent alone is not one of those methods.

This matters in both directions. A spoofed “GPTBot” can inflate a report of AI crawler activity. It can also make it look as if an operator ignored your robots.txt when the requests came from someone else.

If you use WP Visibility’s optional AI Traffic Lens module (off by default in 2.10.3), know what it measures. It counts requests whose user agent contains a known AI crawler name, plus visits referred from assistants, and it stores no IP addresses. Its counts are therefore claimed visits. It also sees only requests that reach WordPress; a page served from a full-page cache or a CDN never runs the plugin. The command that prints its report is in the WP-CLI reference.

From claim to verified visit. Claimed identity: The user agent names a known crawler. Operator method: IP range file or reverse and forward DNS. Label the request: Verified, unverified, or inconsistent.
A user agent only starts the check. Verification depends on the method each operator publishes, and some lists confirm the operator but not the specific bot.

Collect a sanitized request sample

Work from raw access logs: your host’s logs, your web server’s, or your CDN’s. Pick a fixed window, such as seven days.

  1. Find where the real client IP is. If your site sits behind a CDN or proxy, your origin’s logs may show the proxy’s addresses. Use the CDN’s logs, or the request header your CDN documents for the original client IP.
  2. Extract claimed AI crawler requests. For a standard combined log format:
grep -E "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|Googlebot" access.log > claimed.log
awk '{print $1}' claimed.log | sort -u > claimed-ips.txt

Adjust the file names and the field number to your log format.

  1. Keep the working files private. IP addresses in logs can identify people, especially in spoofed requests from ordinary connections. Do not paste raw logs into shared documents or tickets. For anything you share, replace IP addresses with the verification result, or truncate them.
  2. Add a deliberate spoof case. From your own connection, request a page while claiming to be a crawler:
curl -s -o /dev/null -A "Mozilla/5.0 (compatible; GPTBot/1.4; +https://openai.com/gptbot)" https://example.com/

Use your own domain. Your request should appear in the logs with a GPTBot user agent and your IP address. It must come out of the next step as unverified; if it does not, the check is not working.

Apply supported verification methods

Each operator documents its own method.

Operator Published method Source, checked September 27, 2026
OpenAI Separate IP range files for OAI-SearchBot, GPTBot, ChatGPT-User, and OAI-AdsBot OpenAI crawler documentation
Anthropic One IP range file for its bots Anthropic crawler help article
Perplexity Separate IP range files for PerplexityBot and Perplexity-User Perplexity crawler documentation
Google Reverse then forward DNS lookup, or IP range files Verify requests from Google crawlers and fetchers

Anthropic’s article says that if a crawler “has a source IP address on this list, it indicates that the crawler is coming from Anthropic.” Because the article publishes a single list rather than one per bot, an IP match confirms Anthropic but not which of its bots made the request. Perplexity’s documentation notes that WAF changes “may take some time to propagate” and recommends refreshing IP ranges from its endpoints periodically; the same applies to any list you download.

Check IP ranges with a script. This Python script reads IP addresses, one per line, and reports which published range contains each. It uses only the standard library.

import ipaddress, json, sys, urllib.request

SOURCES = {
    "OpenAI OAI-SearchBot": "https://openai.com/searchbot.json",
    "OpenAI GPTBot": "https://openai.com/gptbot.json",
    "OpenAI ChatGPT-User": "https://openai.com/chatgpt-user.json",
    "Anthropic (any of its bots)": "https://claude.com/crawling/bots.json",
    "PerplexityBot": "https://www.perplexity.com/perplexitybot.json",
    "Perplexity-User": "https://www.perplexity.com/perplexity-user.json",
}

def load(url):
    req = urllib.request.Request(url, headers={"User-Agent": "range-check"})
    with urllib.request.urlopen(req, timeout=20) as response:
        data = json.load(response)
    networks = []
    for prefix in data.get("prefixes", []):
        cidr = prefix.get("ipv4Prefix") or prefix.get("ipv6Prefix")
        if cidr:
            networks.append(ipaddress.ip_network(cidr))
    return networks

ranges = {name: load(url) for name, url in SOURCES.items()}

for line in sys.stdin:
    ip = line.strip()
    if not ip:
        continue
    address = ipaddress.ip_address(ip)
    owners = [name for name, nets in ranges.items() if any(address in net for net in nets)]
    print(ip, "->", ", ".join(owners) or "no match")

Run it with python check_ranges.py < claimed-ips.txt. The file URLs are the ones each operator’s documentation lists; if a URL stops working, return to that documentation rather than guessing a replacement.

Check Google with DNS. Google’s page gives four steps: run a reverse DNS lookup on the IP with the host command, confirm the name ends in googlebot.com, google.com, or googleusercontent.com, run a forward lookup on that name, and confirm it returns the original IP.

host <ip-from-your-log>
host <name-returned-above>

Both lookups must agree. A reverse lookup alone can be set by whoever controls the IP’s DNS.

Report verified and uncertain traffic

Match each request’s claimed identity with the verification result:

Claimed user agent Verification result Label
GPTBot IP in OpenAI’s GPTBot range Verified GPTBot
GPTBot IP in no OpenAI range Claimed, unverified
ClaudeBot IP in Anthropic’s range Verified Anthropic (bot not distinguished by IP)
Googlebot Reverse and forward DNS agree on googlebot.com Verified Google
Any AI crawler name IP matches a different operator’s range Inconsistent; investigate
No crawler name Any Unclaimed, outside this report

Your spoof request should be in the second row.

A useful report states:

  • The log source and time window.
  • Verified requests per operator and bot, with status codes. A verified crawler receiving 403 or 429 responses is an access problem; the firewall rule guide traces it.
  • Claimed but unverified requests, as a separate total. Do not add them to the operator’s figures.
  • Anything you could not check, such as requests with no usable client IP.
  • The date the IP lists were downloaded.

Do not conclude that an operator ignored robots.txt from unverified requests. Check the verified group against your rules, allowing for the adjustment periods operators mention, such as OpenAI’s note that, for search results, it can take about 24 hours after a robots.txt update for its systems to adjust (OpenAI crawler documentation, checked September 27, 2026). The policy side is in the AI crawler policy guide.

Checklist

  • Logs contain the real client IP, not a proxy’s.
  • A spoofed request was added and came out unverified.
  • IP ranges downloaded fresh, with the date recorded.
  • Google requests checked with reverse and forward DNS.
  • Verified, unverified, and unclaimed traffic reported separately.
  • No raw visitor IP addresses in anything shared.

Read next

WordPress SEO with your own assistant.

WP Visibility is $99 a year for unlimited sites, client sites included, with a 30-day refund. Use its SEO tools in WordPress or connect a supported assistant. Read how proposal review and permissions work.