Skip to main content
← Back to Radar

Methodology

How Radar gathers AI response data through live queries, calculates your AI Readiness Score, handles LLM response variance, and detects hallucinations. Every metric is transparent and traceable: we store the measured inputs, timestamps, and provider coverage behind each score. LLM responses are non-deterministic, so identical reruns are not guaranteed (see LLM Variance Handling below).

The 13 Audit Dimensions

Each dimension measures a specific aspect of your AI visibility infrastructure.

AI Crawl CheckProposed weight: 9%

Measures: Access for 17 AI, search, and SEO crawler user-agents, with retrieval access and training policy assessed separately.

Why it matters: Retrieval crawlers need access to fetch your content for AI answers. Training access is a separate policy choice.

Robots.txt AnalysisProposed weight: 8%

Measures: Deep parse of robots.txt rules for 20 known AI and search bots.

Why it matters: Misconfigured robots.txt is the most common reason sites are invisible to AI.

llms.txt ValidationProposed weight: 8%

Measures: Whether your llms.txt file exists, is valid, and contains useful structured content.

Why it matters: llms.txt is a proposed file that summarizes a site for language models. The current methodology still scores it, although Google says Google Search ignores it and no major AI engine documents using it to choose citations (checked September 2026). We are reviewing its weight for the next methodology version.

AI Readiness ScoreProposed weight: 8%

Measures: 5-category audit across technical, content, structured data, authority, and discoverability.

Why it matters: A holistic assessment of whether your site is optimized for AI citation.

Citation TrackerProposed weight: 9%

Measures: Whether the ChatGPT, Claude, Gemini, and Perplexity APIs mention and cite your brand.

Why it matters: The ultimate measure: are AI models actually talking about you?

Reddit MonitorProposed weight: 5%

Measures: Brand mentions across Reddit, including detection of seeded or artificial content.

Why it matters: Reddit is a major training data source for LLMs. Your Reddit presence shapes AI outputs.

AEO Page AuditorProposed weight: 8%

Measures: Answer Engine Optimization readiness across 6 categories for a specific page.

Why it matters: Page-level optimization determines whether AI can extract quotable answers.

Citation TesterProposed weight: 8%

Measures: Whether AI platforms cite your specific page when asked a relevant question.

Why it matters: Tests the direct pathway from question to your content.

Source InfluenceProposed weight: 7%

Measures: Which domains are shaping AI narratives in your category.

Why it matters: Understanding the competitive landscape tells you what you are up against.

Prompt SOVProposed weight: 7%

Measures: Your position-weighted share of brand mentions in the AI responses Radar collected for your category.

Why it matters: Shows which brands the measured models name, and how early, when asked about your category. It counts mentions within the sampled responses, not endorsements or market share.

Schema AuditProposed weight: 7%

Measures: Scores contextual JSON-LD coverage and field completeness across 13 supported schema types.

Why it matters: Radar uses these entities as its business-identity and page-context signals.

Hallucination CheckProposed weight: 8%

Measures: Detects factual errors in what AI models claim about your brand.

Why it matters: AI hallucinations can damage your brand. You need to know what is being said wrong.

Brand DisambiguationProposed weight: 8%

Measures: Whether AI engines link your brand name to the right entity or confuse it with a same-named company, person, or product.

Why it matters: If AI describes a different entity that shares your name, buyers get the wrong company. Entity accuracy protects your identity in AI answers.

Scoring Model

The overall Radar score averages the dimensions that completed for your domain, normalized to 0-100. Today all completed dimensions are weighted equally. The proposed weights shown above are our importance ranking and the basis for the calibrated weighting on our roadmap; they do not affect today's score.

A

85-100

Excellent

B

70-84

Good

C

50-69

Average

D

30-49

Poor

F

0-29

Critical

Calibration: The dimension importance ranking started as a judgement informed by the fixes we made on our own site between October 2025 and March 2026: a directional prior from a single case, not a statistically calibrated model, and that case recorded no AI answers before or after the fixes. Calibration is ongoing: calibrated weight sets are built and evaluated against real audit outcomes before any promotion into the live score.

What "good" looks like: A grade summarizes the checks an audit completed. It is not a prediction that AI engines will cite or recommend a site; the live checks in a full audit measure that directly, for the prompts they test.

Scoring Changelog

When a scoring rule changes, a score can move without anything on the site changing. Every such change is dated and explained here, and the same note appears beside the score it affected.

Schema Audit scoring methodology v2 — 3 August 2026

Schema Audit now scores business identity, WebSite, navigation context, and a page type appropriate to the audited content. Article, FAQPage, and SpeakableSpecification are no longer universal requirements, and ProfessionalService, Place, and CreativeWork now have field-completeness rubrics.

Volume now counts typed schema entities rather than type memberships described as JSON-LD blocks, and volume and evaluated-type diversity build one point at a time instead of maxing out at five. If your score changed without a site update, compare runs by scoring-method version; version 1 and version 2 are not directly comparable.

Withheld checks no longer count as zero — 20 July 2026

When a check cannot grade your site — usually because a firewall blocked it — we say so instead of inventing a number. That "withheld" result was still being averaged into your overall score as a zero, which pulled the total down for a check that had explicitly declined to score you.

Withheld checks are now left out of the average entirely. If your overall score went up, nothing on your site changed; the earlier number was unfairly low. We now also show how many checks the score is based on, so a score built on 12 checks is never mistaken for one built on 13.

Scoring methodology updated — 20 July 2026

We fixed how the AI Crawl Check scores a site that returns a normal web page instead of a proper "404 not found" when a file (like llms.txt or robots.txt) is missing. We used to treat that as "couldn't check" and drop it from your score; it actually means the file simply isn't there, so it now counts as missing — the same as any other site without one.

If your score went down, nothing on your site changed for the worse. The old number was too generous, and the new one matches what any site with the same setup would get. The way to raise it is what it always was: add the missing file.

Crawl Integrity Score

A composite metric across AI Crawl Check, Robots.txt Analysis, and llms.txt Validation.

The Crawl Integrity Score averages three component scores: bot access policy for 17 bots (checked with simulated requests from Radar's servers; Google-Extended is a robots.txt token only), robots.txt rule analysis for 20 bots, and llms.txt validation against Radar's rubric. Read the components to see which one moved the average.

LLM Variance Handling

LLM responses are non-deterministic. The same query can produce different results on different runs. Radar addresses this openly.

Methodology

Citation scores query each LLM and display the result. When comparing across runs (LLM Answer Diff), Radar highlights changes explicitly so users can distinguish genuine shifts from model variance.

Why this matters

Most tools in the AI visibility space don't surface non-determinism. Radar acknowledges and handles LLM variance explicitly, turning a vulnerability into a trust signal.

Live Queries vs Snapshot Databases

How Radar gathers AI response data, and why this matters for accuracy.

Live querying beats snapshot databases when AI answers change between runs. Radar runs fresh queries against the ChatGPT, Claude, Gemini, and Perplexity provider APIs in every full audit of a domain. The results reflect what the measured API models returned at audit time, not what they said weeks ago.

What Radar queries

Radar measures responses from the provider APIs (OpenAI gpt-4o-mini, Anthropic Claude Haiku, Perplexity Sonar, and Google Gemini 2.5 Flash). Consumer apps like chatgpt.com or gemini.google.com may differ: they can use different models, memory, personalization, and live web search. Results are a snapshot of the measured API responses at scan time; provider coverage and model versions are shown where available. Where shown in your report, Radar also measures xAI Grok (grok-4.3) as a report-only engine during rollout; report-only engines appear in per-engine results but do not affect scores.

Snapshot approach

Tools build a large prompt library, often hundreds of millions of search-derived prompts, run them in batched cycles, and serve the cached results. Refresh schedules vary by tool and plan, from daily to monthly.

Strength: scale of queries. Weakness: staleness, plus invisibility to AI sub-queries the platform never sees.

Radar live-query approach

Radar issues fresh queries to each measured API model in every full audit. Results reflect those models at audit time, including live search where the model uses it.

Strength: no cached answers, captures current API behavior. Trade-off: smaller per-audit query volume, addressed by repeat audits and explicit variance handling.

97.6%
Underreporting rate found in independent benchmark testing of a leading snapshot-based AI visibility tool against live ChatGPT mentions
Source: independent industry benchmark, 2026

The Dark Query Blind Spot

Most AI retrieval traffic is invisible to prompt-library tools.

When a user asks an AI assistant a question, the model often decomposes that question into multiple internal sub-queries before generating an answer. These sub-queries have no Google search volume, no public footprint, and never appear in keyword-research-based prompt libraries. Industry estimates put this dark-query share at roughly 88% of total AI retrieval traffic.

~88%
Estimated share of AI retrieval traffic generated by internally-decomposed dark queries with zero search volume
Source: industry analysis, 2026

Radar handles this by querying LLMs directly with category and brand-specific prompts that mirror how AI assistants actually retrieve answers, not how Google users phrase searches.

When Snapshots Are Useful

Live querying is not always the right answer.

Snapshot databases earn their keep when you need historical trend lines, fixed comparison surfaces across thousands of brands, or coverage of branded search-volume data that only traditional search engines can supply. Radar focuses on technical AI readiness and live citation behavior, and we recommend pairing live audits with traditional search-volume tools for the full picture. Use the right instrument for the question you are answering.

Hallucination Detection Framework

How Radar identifies and scores factual errors in AI-generated content about your brand.

Ground truth extraction

Radar extracts verifiable claims from your site (meta tags, schema markup, published content) and uses these as the baseline for accuracy comparison.

Claim verification

AI model responses are compared against ground truth. Each discrepancy is flagged with a severity tier: Critical (wrong facts that could cause harm), Major (significant misrepresentations), and Minor (imprecise but not harmful).

Severity scoring

The hallucination score reflects both the number and severity of detected inaccuracies. A domain with zero hallucinations scores 100. Each critical flag reduces the score significantly; minor flags have smaller impact.

Where the Checks Came From

Radar's checks came out of fixes on our own site. That history explains why the checks exist. It is not a validation of the score, because we did not record what AI engines said before and after.

Oct 2025first fixes on our site
515 commits of every kind
Mar 2026Radar platform

Between October 2025 and March 2026 we fixed crawl access, robots.txt rules and structured data on our own site, and the git history records each change. We did not record AI answers before or after, so this is the origin of Radar's checks, not evidence that the score predicts citations.

Read the full origin story →

Related Standalone Tools

Tools outside the 13-dimension Radar audit that complement it.

YouTube Brand Monitor

Google and OpenAI have trained models on YouTube transcripts, but it is not part of Radar's 13-dimension audit. The standalone YouTube Brand Monitor tool scores your YouTube footprint across mention volume, channel diversity, sentiment, reach, and recency using live YouTube Data API queries. Use it alongside Radar for a fuller picture of off-site brand visibility.

Run a YouTube audit →

See it in action

Read a full audit scored by this methodology, or run one on your own domain.