You grepped an access log and found ClaudeBot, GPTBot, Bytespider or another name you did not invite. This index covers 16 AI crawlers and fetchers: what each one does, whether its operator says it honors robots.txt, the exact lines to allow or block it, and what it has actually requested from this site. Every operator claim links the page it was read from.
16 agents from 10 operators · observed traffic logged here since 4 Aug 2026
These agents do four different jobs, and the robots.txt decision is different for each:
Training crawlers (GPTBot, ClaudeBot, Meta-ExternalAgent) collect content for building AI models. Blocking one keeps future models from training on a site; it does not remove the site from any search product.
AI search indexers (OAI-SearchBot, PerplexityBot, Claude-SearchBot, YouBot, Applebot) feed the answer engines' own search layers. Blocking one removes a site from that engine's results, which is the opposite of what anyone measuring AI visibility wants.
User-requested fetchers (ChatGPT-User, Perplexity-User, Claude-User, Meta-ExternalFetcher) fire when a person asks an assistant a question and it opens a page to answer. A hit from one of these IS the AI-visibility event: an assistant just read your page on someone's behalf. Two operators state plainly that these may ignore robots.txt.
Control tokens (Google-Extended, Applebot-Extended) never appear in a log. They only change what already-crawled content may be used for.
The observed column is this site's own logging, not a global number. It is first-party evidence of how these bots behave against a real property, which is exactly what the tools in the index sell at scale.
Get changes to the index
One email when prices, engines or entries change. Nothing else.