Separate robots permission, simulated requests, genuine visits and observed citations using a reproducible verification procedure.

In this article
A permissive robots.txt and an HTTP 200 response do not establish that a genuine AI crawler read your catalog. The file describes permission. A request from your computer measures access from your IP. To establish a real crawler visit, compare trusted request logs, the observed source address and the provider's official network ranges.
This distinction prevents two mistakes: opening a firewall to anyone who inserts a bot name into User-Agent, and reporting a simulated request as evidence of ChatGPT visibility. The following procedure concerns public store or publication pages, while preserving protection for private operations.
Classify crawlers by their actual function
OpenAI separates OAI-SearchBot, GPTBot and ChatGPT-User. Search, training and user-triggered actions are separate functions. Permitting search does not require permitting training. Record those choices before editing robots.txt.
Create an inventory with provider, purpose, robots token, official IP endpoint, current policy and responsible owner. Include a last-checked date. A copied third-party list should not become the permanent operational source.
Google's verification documentation distinguishes crawler and fetcher families. They do not all use the same address file. Associate a function with its appropriate source rather than allowing every network associated with Google.
Check the current 2026 endpoints
Google announced an IP-file location change on March 31, 2026, moving toward /crawling/ipranges/. If a scheduled job still requests /search/apis/ipranges/, examine its final destination and returned JSON. A 200 response alone does not validate the expected schema.
These are three official sources for a search-access inventory:
| Function | Official prefix file |
|---|---|
| Google common crawlers | common-crawlers.json |
| OpenAI search | searchbot.json |
| PerplexityBot | perplexitybot.json |
Our GET checks on October 3, 2026 returned 200 and a prefixes array for all three endpoints. This establishes their reachability and structure at the observation time. It does not establish that any provider visited your website.
Start with a trustworthy observed address
Obtain the source address from a trusted edge or server log. If an origin sees only its proxy's address, document how that proxy supplies the original client IP. A header freely supplied by the visitor is not identity evidence.
For each event, preserve UTC time, host, path, status, firewall action and a request identifier. Avoid including cookies, credentials, emails or personal query values in an SEO worksheet. Match the event with the version of the public page served at that time.
Download the official file with normal certificate validation:
curl --fail --silent --show-error https://openai.com/searchbot.json -o searchbot.json
Our free Python prefix checker uses only the standard library. It examines a local JSON file and does not make network requests or modify security rules:
python3 crawler-ip-check.py searchbot.json 192.0.2.10
The address in this example is reserved for documentation. Replace it with the observed log address. NO_PREFIX_MATCH says the address is outside the loaded ranges; it does not by itself establish malicious behavior or a provider error. PREFIX_MATCH establishes address membership, not complete product identity or subsequent use of the page. Check the claimed function and request context too.
Limit a firewall exception to the demonstrated problem
Perplexity's crawler documentation describes combining bot identification with official IP ranges. Apply that principle to the relevant purpose and public paths. Do not remove protection for customer accounts, administration or payment operations just to make an article accessible.
Before changing a policy, preserve a denied event and its configuration. Afterwards, look for a new authenticated request to the same public page. Compare status, content type and HTML content: a challenge page can sometimes look like a successful response when only status is examined.
Define rollback behavior. If the update job cannot fetch or parse a valid prefix file, report the problem instead of replacing working rules with an empty list. Record the age of the last available snapshot. Temporarily retaining it can be an operational choice, but it is not proof that the list remains complete.
An implementation can stage its changes before activation: parse the downloaded data, reject missing prefixes, validate both IP versions, compare the new count with the previous observation and have a human review an unexpected structural change. A count difference is a review trigger, not a universal threshold proving corruption.
Keep four different outcomes separate
A local request with an AI-crawler User-Agent can reveal some policy differences. It does not originate from that provider's network. Label it as a simulation even when it returns 200.
A verified log event demonstrates a real request to a path at a time. It does not establish inclusion in an index, a citation in an answer or a resulting order. Track verified responses, retrieved pages, observed citations and attributed visits or purchases separately. When a data source is unavailable, mark the result unavailable rather than zero.
For Google, AI-search optimization retains SEO's indexing and quality requirements. A navigation file for other systems cannot replace those requirements. For Q4 commerce, prioritize facts a buyer can inspect: the actual offer, variant, current price, availability and delivery conditions. Our Q4 dossier and free resources organize those checks without promising an AI citation or revenue increase.

