Can AI bots find your content?
Check access policies for 17 crawlers, including OpenAI and Claude search, browsing, and training agents.
Ranking on Google says nothing about AI access.
AI search engines like ChatGPT, Perplexity, and Gemini use their own web crawlers to index content. Each bot has a unique user-agent, and many websites accidentally block them through robots.txt rules written for traditional search.
An AI crawl check audits your site against 17 crawler user-agents, separates retrieval access from training policy, and flags missing structured data and llms.txt files that help AI models understand your content. Without this visibility, your site may be invisible to AI search results even if it ranks well on Google.
Search crawlers and training crawlers do different jobs
| Crawler type | Examples | What it is for | What the vendor says about blocking it |
|---|---|---|---|
| Training | GPTBot, ClaudeBot | Collect content that may be used to train AI models | Anthropic: restricting ClaudeBot signals that your future content should be excluded from its training data. |
| Search | OAI-SearchBot, Claude-SearchBot, PerplexityBot | Surface and link pages in the assistant's search results | OpenAI: sites that disallow OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links. Anthropic: blocking Claude-SearchBot may reduce your visibility in user search results. Perplexity: PerplexityBot is not used to train AI foundation models; its page does not say what blocking does. |
| User-triggered | ChatGPT-User, Claude-User | Fetch a page when a person asks the assistant to | OpenAI: robots.txt rules may not apply to ChatGPT-User, because a person starts the fetch. Anthropic: disabling Claude-User may reduce your visibility for user-directed web search. |
To stay in ChatGPT search answers, make sure your robots.txt does not disallow OAI-SearchBot. An explicit rule looks like this:
User-agent: OAI-SearchBot Allow: /
Sources: OpenAI crawler docs, Anthropic's crawler article and Perplexity's bot docs, checked 26 Sep 2026.
What we check.
4 checks, each with a plain-English reason it matters. Every one maps to a line in the report.
17 Bot User-Agents
HEAD requests as Googlebot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, GPTBot, and more.
robots.txt Analysis
Parse allow/disallow rules per bot, sitemap directive, and crawl-delay.
Structured Data
JSON-LD extraction: Article, FAQPage, Organization, Product, BreadcrumbList.
llms.txt Detection
Check for the emerging standard that helps AI understand your site.
How We Score (And Why It's Different)
Most crawl checkers stop at robots.txt. We score what actually determines whether AI systems cite you: structured data quality, llms.txt presence, and content accessibility. This methodology comes from our own work optimizing for AI search.
Do AI retrieval bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot) have access? Is there a deliberate policy for training crawlers? A smart robots.txt isn't just allow/block. It's a strategy.
JSON-LD is how machines understand your content. We check for page-type-appropriate schema: Organization for homepages, Article for posts, BreadcrumbList for navigation. FAQPage also scores, though Google stopped showing FAQ rich results on May 7, 2026.
A proposed file that summarizes your site for language models. We score structure, entity definitions, URLs, and use policy. It is optional: Google says Google Search ignores it, and no major AI engine documents using it to choose citations.
Can AI crawlers actually read your content without executing JavaScript? We check server-side rendering, title/description length, and framework detection.
Built on our own research
This tool reflects what we learned optimizing pixelmojo.io for AI search. Every scoring criterion comes from real implementation and measurable results, not theory.
Related: the free AI visibility tools guide, the robots.txt Analyzer and the AI Readiness Checker.
Want the full picture across four AI engines?
Radar orchestrates 13 AI visibility audits in staged batches: crawl access, llms.txt, structured data, brand citations, and more. It then surfaces cross-tool conflicts and ships a prioritized action plan for your domain.
Run it on your site. See what AI can read.
Free, no account, result on this page. When you want the whole picture, Radar runs all 13 audits in one pass.