Method 1: official IP lists (recommended)
Most major crawlers publish their IP ranges as JSON. Check whether the request IP is in the list of the bot it claims to be:
| Crawler | List |
|---|---|
| Googlebot | googlebot IP ranges |
| Bingbot | bingbot IP ranges |
| GPTBot | GPTBot IP addresses |
| OAI-SearchBot | OAI-SearchBot IP addresses |
| ChatGPT-User | ChatGPT-User IP addresses |
| PerplexityBot | PerplexityBot IP addresses |
| Applebot | Applebot IP ranges |
import ipaddress, json, urllib.request
def load(url):
data = json.load(urllib.request.urlopen(url))
return [ipaddress.ip_network(p.get("ipv4Prefix") or p.get("ipv6Prefix")) for p in data["prefixes"]]
GOOGLEBOT = load("https://developers.google.com/static/search/apis/ipranges/googlebot.json")
def is_googlebot(ip: str) -> bool:
addr = ipaddress.ip_address(ip)
return any(addr in net for net in GOOGLEBOT)
Refresh the lists daily; providers add ranges without notice.
Method 2: reverse DNS + forward DNS
Google and Bing also document DNS verification:
- Reverse lookup the IP:
host 66.249.66.1→crawl-66-249-66-1.googlebot.com - Check that the name ends in
googlebot.com,google.comorgoogleusercontent.com(Bing:search.msn.com). - Forward lookup the name and make sure it resolves back to the same IP.
DNS verification is slower (two lookups per new IP), so cache the result.
Method 3: let your CDN do it
Cloudflare marks verified bots (cf.client.bot in WAF rules) and lets you allow or challenge them in one rule. Other CDNs offer similar bot management features.
Blocking AI crawlers
Training crawlers like GPTBot respect robots.txt:
User-agent: GPTBot
Disallow: /
Search and user-triggered agents (OAI-SearchBot, ChatGPT-User, Perplexity-User) are separate user agents. Block them only if you do not want to appear in AI search answers.