AI Visibility Numbers Are Unreliable. Measure Them Anyway.
Generative engine optimization (GEO), also called answer engine optimization (AEO), has become a budget line at most consumer brands. The measurement underneath it has not earned that confidence. Repeated runs of the same prompt return different brands and different sources, purchase-intent questions are the least stable query type of all, and published research shows an apparent move from 8% to 11% citation share cannot be distinguished from random noise. Most AI visibility dashboards report a single number to one decimal place anyway. The argument here: build the practice regardless. Not because the number is accurate, but because a shared methodology, a common vocabulary and a team that can reason about probabilistic measurement are worth more right now than the metric itself.
Two things about generative engine optimization are true at the same time. The measurement underneath it is unreliable enough that most of the numbers being reported to boards this year would not survive a statistics review. And the brands waiting for better measurement before they start will be a year or two behind the ones that don't.
That tension is worth working through carefully, because a large amount of money is about to move on the strength of numbers almost nobody can audit.
Generative engine optimization (GEO), also called answer engine optimization (AEO), is the practice of trying to influence whether and how a brand appears when a customer asks an AI assistant a question. The adjacent practice, AI visibility monitoring, tries to measure how often that happens.
The framing matters. GEO is not a channel. It is a layer of influence, and the click from a citation inside ChatGPT is the smallest and least interesting part of it. The relevant question is not how many sales an agent completed. It is how many purchases an agent shaped -- the same distinction that separates agentic commerce from agent-executed commerce, and social commerce from in-platform checkout.
That distinction is also the first clue about why this is so hard to measure.
This is the part executives find hardest, and the one that should reset expectations.
Search is close to deterministic. The same query returns roughly the same ranked list to everyone, and a position earned through good work is a durable asset -- often six months or more of advantage until the algorithm updates. That durability is what made SEO investable.
Generative systems are probabilistic. The answer is composed fresh on every request. There is no position to hold and no state to defend. Give a brand a magic wand today that revealed its exact standing and the precise levers to improve it, and the picture would drift by next week without anyone touching anything -- because the model changed, the retrieval layer changed, a competitor published, or the customer's own history changed.
GEO is closer to running a continuous experiment than to holding a ranking. Teams that budget for it as a project will be disappointed. Teams that budget for it as an ongoing measurement function will not.
Most GEO conversations quietly assume one kind of agent. There are two.
Horizontal agents -- ChatGPT, Gemini, Claude, Perplexity, Copilot -- are trained on the open web and sit closest to discovery. Vertical agents are retailer-owned and catalog-bound -- Amazon's Alexa for Shopping (formerly Rufus), Walmart's Sparky, and the assistants appearing inside grocery, pharmacy and specialty apps -- and they sit closest to the transaction. Alexa for Shopping is trained on Amazon's catalogue and its Prime customer behavior. ChatGPT is not. Visibility in one says very little about visibility in the other.
Almost no tool attempts both. The monitoring category is built around horizontal assistants because those can be prompted from the outside. Retailer agents cannot be queried at scale and return almost no data to brands. So the surface nearest the purchase is the one few brands try to measure, and the surface furthest from it generates all the dashboards.
Two fundamentally different methods, and the distinction rarely appears on a sales slide.
Synthetic prompt monitoring generates a prompt set, runs it against the engines on a schedule, and reports what came back. Repeatable and auditable. Its weakness is that someone had to guess what customers ask.
Real prompt monitoring uses consented consumer panels and observes conversations that actually occurred. One such panel from Measure Protocol, spanning apps, browsers, search and AI assistants, reported in 2025 that more than one in five ChatGPT conversations showed commercial intent. Several GEO vendors now build panels of this kind into their products. Panel data is the closest thing this field has to observed rather than stated preference, and observed preference is almost always worth more.
Neither is sufficient alone. Synthetic testing shows how engines respond to a controlled stimulus. Panels show what people actually ask. The useful programs run both.
One caution applies to every vendor here: very few disclose how the data is collected. Which surface, which model build, logged in or out, which region, how many repetitions per prompt, how variance is handled. A buyer who cannot get those answers in writing is not buying a measurement. They are buying a number.
The fraud grows with them, and the category already contains tactics with no evidence behind them. Ninety-seven percent of published llms.txt files received zero requests in May, and among files that did get traffic the largest single category of requester was SEO audit tools checking whether the file existed, Ahrefs analysis of server logs from 137,000 domains" href="https://ahrefs.com/blog/llmstxt-study/" rel="nofollow noopener" target="_blank">according to an Ahrefs analysis of server logs from 137,000 domains. Google has said the file is not needed. It is still being sold.
Expect worse as spending rises: fabricated benchmarks, content seeded to game retrieval, and proprietary scores only the vendor's own methodology can improve. Expect the pace of change to accelerate rather than settle.
Because in 2026, the deliverable is not the number. It is the organization.
There is real evidence the influence is worth chasing. Users who saw a brand recommended by ChatGPT were 2.5 times more likely to visit that brand's site within seven days, with 56% of that traffic arriving through branded search rather than an AI referral, Similarweb clickstream research" href="https://www.similarweb.com/blog/insights/ai-news/ai-visibility-downstream-impact/" rel="nofollow noopener" target="_blank">according to Similarweb clickstream research. Similarweb labels the finding correlation rather than causation, and the study covers United States desktop users across three verticals. It looks a great deal like measuring a billboard: lift, not clicks.
One caveat belongs on all of it. Nearly every study cited here, on both sides, was produced by a company that sells software in this category. That does not make the findings wrong, but it should shape how confidently they are read. The most rigorous work is academic, and its central conclusion is that everyone else's numbers need error bars.
Write the methodology down before writing the number down, and publish both together. Report ranges rather than point estimates. Build the prompt set from observed customer language -- site search logs, service transcripts, review text, panel data -- not from what the brand team wishes people asked. Track vertical and horizontal agents separately and expect them to disagree. Keep visibility and business outcomes in separate columns, then measure incrementality the only way that has ever worked, through holdouts and matched-market tests. Ask every vendor, in writing, how the data is collected. And do the unglamorous work first: clean product data, content that answers real customer questions, and a firewall that is not quietly blocking the crawlers.
The measurement will be wrong, in ways that can be enumerated. That is the useful kind of wrong. The capability being built -- a team that can reason about probabilistic, agent-mediated customer behavior -- will matter for considerably more than chatbot mentions.
Just don’t put three decimal places on it.
