What Can You Actually Control in GEO?
Generative engine optimization (GEO) is the work of making a site easy for AI assistants to reach, understand and quote, and of measuring what they say about you. As of October 2026 you control four things: which AI crawlers can reach your pages, how quotable and accurate those pages are, how clearly your business is identified, and how you measure. You do not control what an engine writes.
That last sentence is the one most GEO advice leaves out. An earlier version of this playbook promised lift percentages for specific formatting tactics. Those numbers came from a lab benchmark, and the wider research since then has not found a stable effect. This version keeps the steps that hold up and says plainly where the evidence stops.
TL;DR
- Let the right AI crawlers in. Search and user-triggered crawlers decide whether engines can fetch you; training crawlers are a content-use choice.
- Check what your server does, not just what robots.txt says. A firewall can block a bot your rules allow.
- Write pages people can quote: answer first, sourced numbers, tables for comparisons. It is an editorial standard, not a guaranteed lift.
- Keep the technical layer honest: structured data that matches the page, accurate sitemap dates, content in the HTML the server sends.
- Make your business unambiguous, especially if another business shares your name.
- Measure referral traffic and mentions separately. On our own site they ranked the four engines almost in reverse.
GEO is mostly access, accuracy and measurement. Do those well, measure clicks and mentions as two different things, and treat any promised citation lift with suspicion.
The GEO playbook in five steps
As of October 2026
Access
the right crawlers can fetch you
Quotable pages
answer first, sources shown
Technical layer
accurate markup, honest dates
Clear entity
no confusion with namesakes
Measurement
clicks and mentions, separately
Step 1: Which AI Crawlers Should You Let In?
Let in the crawlers that fetch pages for AI search and for users, and decide on training crawlers as a separate content-use choice. Each AI platform runs more than one bot, and each one can be allowed or blocked on its own in robots.txt.
| Crawler type | Examples | What blocking it does |
|---|---|---|
| Training | GPTBot, ClaudeBot, Google-Extended, CCBot, Applebot-Extended | Keeps your content out of model training; a content-use choice |
| Search | OAI-SearchBot, Claude-SearchBot, PerplexityBot | Can keep you out of AI search answers |
| User-triggered | ChatGPT-User, Claude-User, Perplexity-User | Stops the assistant fetching your page when a user asks it to |
The vendors say this in their own documentation. In their crawler docs (checked 26 Sep 2026), OpenAI says sites that disallow OAI-SearchBot will not be shown in ChatGPT search answers, and Anthropic says blocking Claude-SearchBot may reduce your visibility in user search results.
Our own policy, as of 2 October 2026, allows the search and user-triggered crawlers of OpenAI, Anthropic and Perplexity, plus GoogleOther and TavilyBot. We also allow the training crawlers GPTBot and ClaudeBot, and since May 2026 CCBot, Google-Extended and Applebot-Extended as well. We block cohere-ai, Meta-ExternalAgent, FacebookBot, Bytespider, Diffbot, Omgili, Omgilibot and img2dataset. Our dated note on these policy changes explains how the policy moved, and why we no longer claim it changed our citations.
Why check the server and not just the rules?
Because robots.txt is a request, and your server decides what actually happens. A content delivery network or firewall can block a bot that your rules allow, and a page can return content that only appears after JavaScript runs. The AI Crawl Checker fetches your page as each bot and shows the result, which is the quickest way to see whether the rules and the server agree.
Step 2: How Do You Write Pages People Can Quote?
Write each section so its first one or two sentences answer the question on their own, put the source next to every number, and use a table for any comparison. These habits make a page easier to read, scan and check. They are an editorial standard, not a proven way to raise AI citations.
The lift percentages that circulate for these tactics, such as 30 to 40 percent for adding statistics, come from the Princeton and Georgia Tech GEO paper, which measured them on a benchmark built for the study. A July 2026 survey of 45 GEO studies found no stable effect across platforms. Google says in its May 2026 guide to its generative AI features that those features need no special formatting.
What we do on our own posts, and recommend:
- Answer first. The opening sentences of each section should stand alone as an answer.
- Show the source. Put the source and date beside any number, so a reader can check it.
- Use tables for comparisons. Comparisons in prose are hard to scan and easy to misread.
- Phrase headings as the questions people ask. It helps readers find the section they need.
- Date anything that can change. Prices, features and platform behavior go stale.
Step 3: Which Technical Signals Still Matter?
Three technical signals still matter, as of October 2026: content that is in the HTML your server sends, structured data that matches the visible page, and sitemap dates that are accurate. None of them is an AI-specific trick. They are the basics that let any crawler read your site correctly.
Server-rendered content. Vercel's analysis found that the major AI crawlers fetched some JavaScript files but did not execute them, so content that only appears after client-side rendering may be invisible to them. Put the content that matters in the HTML.
Structured data that matches the page. Accurate Organization, Article and Product markup helps search engines understand a page. Google says no special structured data is needed to appear in its AI features (checked 26 Sep 2026), so treat schema as accuracy, not as a lever. Remove anything that does not match the visible page, such as ratings with no real reviews behind them.
Honest sitemap dates. Google's sitemap guide says it uses the lastmod value "if it's consistently and verifiably" accurate (Google Search Central, updated 8 July 2026). Moving dates forward without a real change teaches crawlers to ignore them.
llms.txt is optional. Google says Google Search ignores it, and no major AI engine documents using it to choose citations (checked September 2026). An SE Ranking study of about 300,000 domains found no correlation between llms.txt and AI citations. Some coding agents do fetch it, so it can be worth a few minutes, but not more.
Step 4: Is Your Business Unambiguous?
Make sure an engine cannot confuse you with another business. Use one consistent name, an About page that says what you do and where, and Organization markup that links to your real profiles. If another business shares your name, say what distinguishes you in plain words on the pages engines read.
This matters more than it sounds. In July 2026 an audit of one of our customers found three of four engines describing a different business with the same name. We wrote up what happened, and what we rebuilt in Radar afterward, in What a Wrong-Company Audit Taught Us About AI Visibility. The Brand Disambiguation check looks for exactly this.
Step 5: How Do You Measure GEO Without Fooling Yourself?
Measure two things separately: referral traffic, which counts people who clicked through from an AI assistant, and mentions, which count how often sampled answers name you. They measure different events. On our own site in the first half of 2026, they ranked the four engines almost in reverse.
This section merges our June 2026 post "Our Biggest AI Referrer Cites Us the Least". Its numbers were rechecked against the source data on 2 October 2026, and its wording is corrected: the scores it called citation scores are mention rates from sampled answers.
AI referral sessions to pixelmojo.io
GA4, 22 March to 19 June 2026, all source and medium rows for each assistant
| Engine | Referral sessions (GA4) | Average mention rate (monitor) | Runs where it never named us |
|---|---|---|---|
| Claude | 166 | 11.6% | 162 of 290 |
| ChatGPT | 56 | 29.0% | 0 of 290 |
| Gemini | 46 | 49.6% | 0 of 290 |
| Perplexity | 38 | 50.0% | 0 of 290 |
The mention rates come from our own citation monitor, which asks each engine's API a fixed set of tracked prompts and records how often the answers name Pixelmojo. Between 30 March and 20 June 2026 it completed 290 runs with all four engines. Claude's answers named us in none of the tracked prompts in 162 of those runs. ChatGPT, Gemini and Perplexity each named us at least once in every run.
Why do the two rankings disagree?
Because a click and a mention are different events. A referral happens when a person decides to click a link in an answer. A mention happens when the engine names you, whether anyone clicks or not. An engine can send real visitors through a few answers while naming you rarely across many prompts, and the reverse.
There are two more limits worth stating. Our monitor sampled each engine through its API, and the answers a person sees in the consumer app can differ, because the apps may browse the web and add links. And a mention rate on our tracked prompts is not a share of all conversations about our category. Neither number is the whole picture, which is why we keep both.
Two views of the same AI presence
What each measurement can and cannot see
- Counts people who clicked a link in an answer
- Blind to answers that ended without a click
- Cannot tell an accurate mention from a wrong one
- Counts how often answers name you, on prompts you choose
- Blind to people who never ask those prompts
- API answers can differ from what app users see
How do you set up AI referral tracking in GA4?
Create a custom channel group in GA4 with an "AI Traffic" channel whose source matches the AI assistants, for example chatgpt.com|perplexity.ai|claude.ai|gemini.google.com|copilot.com, and place it above Referral so those sessions are not counted as generic referrals. Then report AI sessions by source over a fixed window, the way the table above does.
How do you track mentions?
Pick a fixed set of prompts your buyers would ask, run them across the engines on a schedule, and record for each answer whether it named you, whether it linked you and whether it described you correctly. Keep the counts, not just the rates. How to Track AI Citations walks through it, and a paid Radar audit runs the prompts for you across ChatGPT, Claude, Gemini and Perplexity.
What About Perplexity Specifically?
Perplexity answers from a live web search and lists its sources, so the basics of search matter more there than anywhere else. In our own monitor it had the highest mention rate of the four engines, 50.0%, and named us in every run.
- Let PerplexityBot and Perplexity-User in. A blocked bot or a page that only renders after JavaScript is the most common reason a live-search engine skips a page.
- Rank in ordinary search. Pages that already rank are the ones a live search tends to find.
- Keep pages dated and current. A visible update date and current figures help for anything time-sensitive.
A 30-Day GEO Plan
A month is enough to fix the basics and set up honest measurement. The plan below is our recommendation, not a promised outcome.
| Week | Focus | What to do |
|---|---|---|
| 1 | Access | Decide your crawler policy by type, update robots.txt, and check the server with the AI Crawl Checker |
| 2 | Entity | Make your name, About page and Organization markup consistent; check for same-named businesses |
| 3 | Content | Rewrite your five most important pages answer first, with sources and dates |
| 4 | Measurement | Add an AI channel in GA4 and run a fixed prompt set across four engines; keep the counts |
After the month, recheck on a schedule. A recheck shows what changed, not why: the engines change on their own, so a better result after a fix is encouraging, not proof.
What Changed in This Post?
We rewrote this post on 2 October 2026 and merged in our June 2026 post "Our Biggest AI Referrer Cites Us the Least", whose old address now redirects here. The changes:
- lift percentages for formatting tactics are removed; they came from a lab benchmark and are not a general effect;
- third-party case study figures, an AI referral volume figure and an outdated monitoring tool price are removed;
- the referral and mention comparison is kept, rechecked against GA4 and our monitor data, and its terms are corrected: mention rates, not citation scores or recommendations;
- the crawler section now matches our current robots.txt policy and the vendors' documentation.
What This Series Covers
This is Part 3 of The AI Search Playbook. Part 2 separates SEO, AEO and GEO. Part 4 is a dated note on what we changed on our own site in February 2026.
GEO Playbook: Questions Readers Ask
Common questions about this topic, answered.
