What Should an AI Visibility Report Prove Before You Act on It?
An AI visibility report should prove what was measured, how it was measured, and what the result can and cannot support. Most reports show a score or a share. Few show the evidence underneath, and without that evidence a number cannot tell you whether to change your content, your budget, or your agency.
AI visibility is several outcomes, not one. An AI answer can name your brand, recommend it for a need, cite a page on your site as a source, send a visitor, or influence a purchase. Each of those needs different evidence. Treating them as interchangeable is how a report sends a strategy in the wrong direction.
This article proposes a standard for that evidence. We call it the AI Visibility Evidence Standard, version 1. It is our proposal, not an established industry standard. We apply it to a real Radar report of our own site, and we publish where our own product falls short of it.
TL;DR
- AI visibility includes at least five different outcomes: a mention, a recommendation, a cited source, a visit, and a sale. A useful report keeps them apart.
- In a September 2026 audit of product queries, the domains ChatGPT and Gemini showed through their APIs overlapped only 12.0 and 14.8 percent with their own consumer apps. A report must say which channel it measured.
- Semrush found that 61.7 percent of 3,981 AI domain appearances were citations with no brand mention. Counting citations as brand visibility overstates it.
- Our proposed standard asks every report to disclose eight things: channel, prompt set, observations and uncertainty, definitions, denominators, source evidence, comparability, and outcomes.
- On a real Radar audit of pixelmojo.io, our brand was named only for prompts that contained our name, and in 0 of 8 unbranded category answers. The headline consistency line hid that split.
- Across five audits of our own site in 22 hours, site-only checks never moved while AI-response checks moved by up to 60 points, and our own code changed in between. Because the method label did not change when the rubric changed, no one can cleanly say why.
- In the same audit, the ChatGPT API named us but described a digital marketing agency, from knowledge it dated to October 2023, and still counted as a mention. A mention says the name appeared, not that the answer was right.
- Judged by what a customer’s report actually shows, Radar fully discloses denominators today. Much of the rest is saved behind the report or not captured yet. We publish the full table.
Before you act on an AI visibility number, ask which channel, prompts, and definitions produced it, and whether it can be compared with the last one. If the report cannot answer, the number cannot support a decision.
A note on evidence. Every Radar figure below comes from production records for pixelmojo.io, read on 25 September 2026. Platform statements are dated and linked to the platform's own documentation or to the research itself.
Why Do AI Visibility Numbers Need an Evidence Standard Now?
Because in 2026 the ways to measure AI visibility multiplied, and they do not measure the same thing. Platforms now publish their own AI reports, research shows that the channel you query changes the answer, and the evidence for popular optimization tactics is thinner than the marketing around it.
Platforms now publish their own AI numbers, and they differ
Google Search Console has a generative AI performance report covering AI Overviews and AI Mode. As of September 2026, its help page lists impressions by page, country, device, and date, excludes Search Labs experiments, and does not list clicks or click-through rate.
Bing's AI Performance report counts citations across Microsoft Copilot, Bing, and select partner experiences. Its Citation Share metric, added on 16 June 2026, is your site's citations as a share of all citations shown for the same grounding query. Bing describes it as an observational metric, not a ranking system or a competitive scoreboard.
A Google impression and a Bing citation are different events, measured on different surfaces. A report that adds them together, or draws them on one trend line, is mixing units.
The channel you query changes the answer
A September 2026 preprint by Uberti-Bona Marin and colleagues evaluated 1,536 responses to product queries from ChatGPT (app and API), Gemini (app and API), and Google AI Overviews. The domains shown through each API overlapped on average only 12.0 percent (ChatGPT) and 14.8 percent (Gemini) with those shown in the matching consumer app. The ChatGPT and Gemini apps shared only 5.4 percent of domains for the same query, and the products they recommended often changed across repeated requests.
The study covers product-recommendation queries, so the exact gap will differ by topic. For product-recommendation queries in that audit, an API answer was not a stand-in for what a customer saw in the app. Many monitoring tools, Radar included, query provider APIs because apps cannot be sampled reliably at scale. That is a defensible choice, as long as the report says so.
A citation is not a mention
In a June 2026 Semrush study of 3,981 domain appearances, 61.7 percent were citations where the brand was never named, 13.2 percent were both a citation and a mention, and 25.1 percent were mentions with no citation. The denominator is every recorded appearance.
A site can supply the evidence for an answer while a competitor gets the recommendation. A report that counts citations as brand visibility, or mentions as endorsement, will overstate what happened.
Optimization claims outrun the evidence
A July 2026 critical survey of 45 studies on generative engine optimization reached a blunt conclusion:
Google's own optimization guide, last updated 10 July 2026, says Google Search ignores llms.txt and similar AI text files, needs no special schema for its AI features, and requires a page to be indexed and eligible for a snippet. If the tactics are uncertain, measurement is how a business learns what works for it. That raises the bar for the reports themselves.
| What changed in 2026 | Evidence | What a report must now do |
|---|---|---|
| Platforms publish native AI reports | Google: AI Overview and AI Mode impressions. Bing: citations and Citation Share | Keep each platform on its own line, with its own definition |
| API and app answers diverge | 12.0% and 14.8% mean domain overlap (product queries) | State API or app, the model, and whether web search was on |
| Citations often lack a brand mention | 61.7% of 3,981 appearances | Count mentions, citations, and recommendations separately |
| Causal evidence for tactics is weak | Survey of 45 studies, July 2026 | Show uncertainty; draw no causal claim from two snapshots |
| Google names technical requirements | Indexed, snippet-eligible, included in Search Console | Report eligibility apart from AI mentions |
What Is the AI Visibility Evidence Standard?
The AI Visibility Evidence Standard is a list of eight disclosures that let a reader judge whether an AI visibility result is meaningful, comparable, and actionable. It is Pixelmojo's proposed standard, version 1, published in September 2026. It does not prescribe a score. It prescribes what any score must show.
| # | Requirement | The report must disclose | Why it matters |
|---|---|---|---|
| 1 | Channel | Provider API or consumer app; the exact model; whether web search or grounding was on | App and API answers can differ sharply for the same question |
| 2 | Prompt set | The exact prompts; which ones contain the brand name; the locale and how it was applied | A branded prompt measures recognition; an unbranded one measures consideration |
| 3 | Observations and uncertainty | Timestamp; usable responses out of attempted; runs per prompt; ranges where judgment is involved | One answer is a sample, not a measurement |
| 4 | Definitions | What counts as a mention, a linked citation, and a recommendation, each counted separately | They are different events with different commercial value |
| 5 | Denominators | Every rate with its count, such as "3 of 8 responses" | A percentage is unreadable without its base |
| 6 | Source evidence | The cited URLs, and whether each one resolves to your site | Whose page was cited decides who earned the credit |
| 7 | Comparability | The method version; a dated log of method changes and whether past results were recomputed or kept; which comparisons are valid; unknown shown as unknown, never as zero | Scores from different methods cannot share a trend line |
| 8 | Outcomes, kept separate | Visits, leads, and revenue apart from exposure; native platform numbers on their own lines; what the report cannot show | Exposure is not a sale, and incrementality is rarely measured |
The difference is easiest to see on one real result, reported two ways. Both columns below describe the same audit, which the next section walks through.
The same audit, reported two ways
The saved record of one Radar Citation Check for pixelmojo.io, 25 September 2026
- No channel, model, or web search setting stated
- Branded and unbranded prompts pooled into one number
- Mentions reported as visibility
- No date, run count, or usable-response count
- Provider APIs, exact models, and web search per engine stated
- Branded, category, and reputation prompts reported apart
- Mentions and linked citations counted separately
- Timestamped, 16 of 16 usable responses, one run per prompt
How Does a Real Report Measure Up? Our Own Site, Annotated
The clearest way to test a standard is on a real report. Below is Radar's Citation Check for pixelmojo.io, run on 25 September 2026 at 01:11 UTC, annotated against the eight requirements from the saved record behind the report. The report a customer sees shows less than that record; the self-assessment later in this article lists exactly what. We chose our own site because we can publish every detail.
The audit asked four prompts in the category "AI visibility audit platforms": one branded ("What is Pixelmojo and what do they specialize in?"), two unbranded category prompts ("What are the best AI visibility audit platforms companies?" and "Compare the top AI visibility audit platforms options. What are the pros and cons of each?"), and one reputation prompt ("What do people say about Pixelmojo? Is it a good company?"). The prompts come from templates with the category inserted, which is why one reads awkwardly. A report should show them exactly as sent.
The record: Radar Citation Check, pixelmojo.io
25 September 2026, 01:11 UTC. Four scored engines plus one report-only engine, one run per prompt.
Prompts: 1 branded, 2 unbranded, 1 reputation
Usable responses from the four scored engines
Unbranded category answers that named Pixelmojo
Citation Check score (grade D)
| Engine and model | Web search | Branded prompt | Unbranded prompts (2) | Reputation prompt |
|---|---|---|---|---|
| ChatGPT, gpt-4o-mini (API) | Off | Named, no link | Not named | Not named |
| Claude, claude-haiku-4-5-20251001 (API) | Off | Not named | Not named | Not named |
| Perplexity, sonar (API) | On | Named, 7 links to our pages | Not named | Named, 2 links to our pages |
| Gemini, gemini-2.5-flash (API, grounded) | On | Named, sources stored as domain only | Not named | Named, sources stored as domain only |
| Grok, grok-4.3 (report-only, not scored) | Off | Named, no link | Not named | Named, no link |
Requirement 1, channel: disclosed, but not in the report itself
Radar's methodology page names the provider APIs and the exact models and states that consumer apps may differ. The report a customer sees shows the engine names, not the model or web search setting for each result. We took both from the saved record. That is a gap under requirement 1.
Web search changes how to read the citation column. Two engines had it on and two had it off. Radar never credits a citation to an engine queried without web search, even if its answer text contains a URL, because that URL was not retrieved as a source for the answer. Their empty citation cells are empty by rule, not by merit.
Requirements 2 and 5, prompts and denominators: the headline hides the split
The report's consistency line reads "Brand mentioned by 3/4 active providers." That is true, and on its own it misleads. All three mentions came from prompts that contained our name. In the two unbranded category prompts, the ones a buyer types before knowing who we are, the four scored engines named us 0 times in 8 answers.
Radar reports that line too ("Mentioned in 0/8 competitive queries"), and it is the line that should lead. A report that pools branded and unbranded prompts is measuring recognition and calling it visibility.
Requirement 4, definitions: a mention is not a correct answer
Here is one prompt from the saved record, answered by two engines in the same audit.
| Field | Perplexity | ChatGPT |
|---|---|---|
| Model and channel | sonar, provider API | gpt-4o-mini, provider API |
| Web search | On | Off |
| Prompt | What is Pixelmojo and what do they specialize in? | Same prompt |
| Checked at | 25 September 2026, 01:11:53 UTC | Same audit |
| Answer excerpt, as saved (formatting removed) | "Pixelmojo is an AI product studio based in Makati, Philippines, founded in 2024 by Lloyd Pilapil.[1][3][4] It builds production AI systems for B2B compani..." | "As of my last update in October 2023, Pixelmojo is a digital marketing agency that specializes in various aspects of online marketing, including search engine optimization (SEO), pay-per-click (PPC)..." |
| Sources captured | 7 URLs on www.pixelmojo.io (listed below) | None |
| Radar classification | Mentioned, primary; linked citation; tone neutral | Mentioned, primary; no citation; tone neutral |
Perplexity's seven captured sources: https://www.pixelmojo.io/press, https://www.pixelmojo.io/about, https://www.pixelmojo.io/, https://www.pixelmojo.io/llms.txt, https://www.pixelmojo.io/blogs/knowledge-graph-llm-visibility-real-data, https://www.pixelmojo.io/blogs/why-we-blocked-ai-training-bots-and-citations-went-up, and https://www.pixelmojo.io/terms-of-service.
Both answers count as a mention, and both were classified primary and neutral. Only one describes the business correctly. The ChatGPT API, without web search, described a digital marketing agency, which is not what Pixelmojo is, from knowledge it dated to October 2023. A report that counts mentions without checking what was said would record that as a win. Radar checks accuracy in a separate Hallucination Check, which is why requirement 4 asks a report to define its terms: a mention says the name appeared, not that the answer was right.
The saved excerpt is also short, about 200 characters around the first mention rather than the whole answer, which limits what anyone can inspect afterwards.
Requirement 6, sources: citations are uneven by engine
Perplexity linked seven of our pages for the branded prompt, including /llms.txt and /terms-of-service, which says something about what it retrieved. Gemini's grounded sources were stored as the domain name only, six entries reading "pixelmojo.io", so the report cannot show which page Gemini used. Under requirement 6 that is a partial disclosure, and we count it as a gap.
Since 25 September 2026, Radar credits a citation only when the link's host is the audited domain or a real subdomain of it, never because the domain appears somewhere in the URL. The same release made the dashboard report new mentions and new linked citations as separate observations.
Requirements 5 and 7, a zero with its denominator
The same audit's share-of-voice check did not find us at all. Its summary reads: "Pixelmojo was not detected in the responses we collected from 4 usable providers (4 attempted) for "AI visibility audit platforms" queries: 20 usable responses in total. 85 other brands were found."
That is how a zero should read. Not "invisible to AI", but absent from 20 specific responses, with a count of what appeared instead.
Requirement 3, uncertainty: a range where judgment is involved
Two checks in the same audit need judgment rather than a rule. For those, Radar asks two different models five times each, claude-sonnet-4-6 and gemini-2.5-flash, and stores the range instead of a single number. On this run the ranges were 52 to 75 and 19 to 39, and both were flagged borderline because the models disagreed. Those judged results sit beside the score; they are not added into it.
What Happens When Method Labels Do Not Track Method Changes?
You lose the ability to say why a score moved. We learned this on our own site: across five full audits in 22 hours, our overall Radar score read 75, 69, 74, 72, and 74, and we cannot cleanly attribute the changes.
The audits ran between 03:00 UTC on 24 September and 01:13 UTC on 25 September 2026. The five checks that read only our site returned the same score every time: Crawl Check 98, AI Readiness 100, llms.txt 100, Robots.txt 82, and Schema 95. The checks that ask AI models moved. The Hallucination Check read 100, 40, 100, 80, and 100. The Answer Engine check read 74, 69, 68, 73, and 74. The Citation Check read 51, 45, 46, 44, and 45.
That looks like model variability, and some of it may be. But the same day, two changes to the Hallucination Check itself were merged, at 06:09 and 09:08 UTC, each within minutes of an audit. We also published content changes to the site. The model answers, our site, and our grader all changed inside one window. Each saved score does carry a method label, but it read the same value in all five audits and through every earlier rubric change, so it cannot separate those three causes.
Radar does keep a public scoring changelog, and its three entries date to July and August 2026. The two Hallucination Check changes of 24 September are not in it. That is a gap in our own practice, and it is the gap requirement 7 is written to close.
Hallucination Check, pixelmojo.io, five audits in 22 hours
Two changes to the check itself were merged the same day, so model, site, and grader effects cannot be separated.
A weekly stability check makes the cause findable
From 17 August to 21 September 2026, Radar re-scored a fixed set of 12 public domains each week with its deterministic readiness check and compared each result with an expected range. Of those 72 results, 70 stayed in range. We traced the two exceptions, both in the week of 14 September: one came from a change in our own robots.txt parser, the other from a change on the site itself. Neither came from the models, and because that check has no model input, neither could have.
That is requirement 7 in practice: a known method, a known baseline, and a record of what changed. It is also the part of Radar's own history we most wanted to get right before asking anyone else to meet the standard.
Where Does Radar Meet the Standard, and Where Does It Fall Short?
We judged each requirement by one question: does the report a customer sees disclose it? A limit the report states counts as disclosure. Evidence that is saved but not shown does not. On that test, Radar's report fully discloses denominators today, and every other requirement is partly shown, partly saved behind the report, or not captured yet. We publish the gaps because a standard its author does not apply to itself is marketing.
| Requirement | Shown in the customer report | Saved or disclosed elsewhere | Not yet captured or shown |
|---|---|---|---|
| 1. Channel | Engine names | Provider API channel and exact models on the methodology page; the model for each result in the saved record | Model and web search setting beside each result; consumer-app observations |
| 2. Prompt set | Every exact prompt, tagged brand, competitive, or reputation | Locale, applied as an instruction to the model | A prompt version on every result |
| 3. Observations and uncertainty | Response coverage when a run is incomplete; a withheld score when coverage is too thin | Timestamp and usable-versus-attempted counts for every engine; ranges for judged checks | Repeated runs of each prompt; the timestamp in the report |
| 4. Definitions | Mention and linked citation as separate results for answers that name the brand, with prominence and tone labels | A citation in an answer that does not name the brand, counted in the URL citation score | That citation shown at answer level; recommendation as its own measure; mention accuracy, which a separate Hallucination Check covers |
| 5. Denominators | Counts beside rates, such as 3/4 and 0/8 | None needed | None |
| 6. Source evidence | Whether each answer that names the brand also linked the site | Cited URLs, matched by host; an excerpt of about 200 characters around the first mention | URLs and excerpts in the report; page-level Gemini sources; full answer text |
| 7. Comparability | Withheld scores shown as withheld, never as zero | A dated scoring changelog on the methodology page; a method label on every saved score | A label that changes when the rubric does; changelog entries for every change to a check |
| 8. Outcomes, kept separate | Exposure only, with no revenue claim | AI referral tracking through to purchase for our own site; an optional Search Console connection | Bing data; incrementality |
Our methodology page already notes that part of Radar's scoring is under review. We are not announcing dates here. The point of the table is that you can check it against any future release, the same way you can check any provider's.
How Can You Apply the Standard to the Report You Already Get?
Ask your provider eight questions, one per requirement, and judge the report by how many it can answer from what it already shows.
- Channel. Did you query the provider APIs or the consumer apps, which models, and with web search on or off?
- Prompt set. Which exact prompts ran, and which of them contained our brand name?
- Observations. When did each run happen, how many responses came back usable, and how many times did each prompt run?
- Definitions. What counts as a mention, a citation, and a recommendation in this report?
- Denominators. What is each percentage out of?
- Sources. Which URLs were cited, and which of them are ours?
- Comparability. Did the method change since the last report, is that change logged with a date, were past results recomputed, and is each zero a real zero or a missing measurement?
- Outcomes. Which numbers are exposure, which are visits or revenue, and which come from Google's or Bing's own reports?
A report that answers all eight from what it already shows is much closer to decision-grade. A report that answers three is the start of a conversation. If you are choosing a provider rather than auditing one, our guide to hiring a GEO agency covers the commercial side. For why AI answers vary at all, see No Brand Controls Its AI Recommendations, and for tracking setup, our guide to tracking AI citations.
That sentence is a useful test of any report. Some eligibility and reporting conditions are visible only inside a platform's own tools. A report that claims to measure them from the outside should say which ones it cannot see.
How Can Other Providers Adopt the Standard?
Any provider can use the AI Visibility Evidence Standard, v1, without asking. Cite it as "Pixelmojo AI Visibility Evidence Standard, v1 (September 2026)" and publish your own table of where you meet each requirement.
We expect it to change. Version 1 does not yet say how many repetitions are enough, how to sample consumer apps reliably, or how to measure recommendations rather than mentions. Those are open questions, and we would rather answer them in public with the rest of the field than set them alone. Critiques and counterexamples are welcome at founders@pixelmojo.io.
How Does Radar Help?
Radar's Citation Tracker runs across four scored engines. Its report shows each prompt, each engine's mention result with a linked-citation marker when the answer also names the brand, and the counts behind every rate. An answer that links the site without naming the brand is counted in the URL citation score but not shown at answer level. The saved record behind it holds more than the report shows today, including the model, timestamp, answer excerpt, and captured cited URLs for each result; those differences are gaps in the table above. For how the scores themselves are built, see A Score You Can Defend.
The AI Visibility Evidence Standard: Questions Readers Ask
Common questions about this topic, answered.
The Standard Is the Point
AI visibility will keep changing faster than any one tactic. What will separate useful work from noise is whether the numbers behind it can be checked. The AI Visibility Evidence Standard, v1, is our proposal for what that checking requires: the channel, the prompts, the observations, the definitions, the denominators, the sources, the method, and the outcomes, each stated plainly.
We have graded our own product against it and published the result. We invite every provider, agency, and in-house team to do the same.
Want to see your own report measured this way?
- Run the Citation Tracker - Every prompt and engine, with counts beside every rate
- Read the Radar methodology - What Radar measures and how
- Contact Us - Talk through your current AI visibility reporting