Why a Single AI Visibility Score Is a Vanity Metric
An AI visibility score is a single number meant to summarise how often AI assistants surface your brand. The problem is not the arithmetic — it is that one number hides the two things that decide whether you get the sale. Being named is not being recommended, and one run is not a measurement. Here is what to demand from any score before it goes in a board deck, including ours.
Viren Inaniyan · September 10, 2026 · AEO-GEO
An AI visibility score is a single number meant to summarise how often AI assistants surface your brand. The problem is not that the number is hard to calculate — it is that a single number hides the two things that decide whether you get the sale. Being named is not being recommended, and one run is not a measurement. This is what to demand from any score before you put it in a board deck, including ours.
A marketing lead showed us a dashboard last month with a number on it: 62. Their AI visibility score, up four points on the quarter. They wanted to know whether 62 was good.
We could not tell them. Not because we did not want to — because the number, on its own, does not carry enough information to answer the question. Two tools can both report 62 while measuring entirely different things, and neither of them has to be lying.
That is worth sitting with, because the score is fast becoming the standard unit of reporting for AI visibility. If it cannot survive a straightforward "is this good?", it is not ready to be the thing your quarter is graded on.
What the number is actually made of
Most AI visibility scores blend some combination of: how often your brand appears in assistant answers, how prominently, across some set of prompts, on some set of engines, over some window. Every one of those "somes" is a decision made by the vendor, and almost none of them get published.
Change the prompt set and the score moves. Change the engine mix and it moves. Change how you weight a passing mention against a direct recommendation and it moves a lot. None of those changes have anything to do with your brand's actual standing.
Practitioners have got blunt about this. One, in a thread titled "I've stopped trusting any AI visibility number", laid out what a score needs to ship with: the exact prompt set and how it is categorised, the number of runs per prompt, the models included, how mentions and recommendations are weighted, and the raw results behind the number. Their conclusion is the line we would put on the wall:
"The methodology is the metric."
Failure one: mentioned and recommended are not the same thing
This is the gap that costs deals, and it is the one a blended score is best at hiding.
An assistant naming your brand in a list of five options is a mention. An assistant telling the shopper to buy you is a recommendation. Both increment most visibility scores. Only one of them sells anything.
A practitioner who ran their own measurement across 76 companies and roughly 1,140 prompts reported companies being named in 44% of answers on average, but recommended in only 10%. Treat those specific figures as one operator's self-reported run rather than an industry constant — but the shape of the finding matches what we see, and the implication is hard to argue with. A brand that is mentioned constantly and recommended rarely looks healthy on a dashboard while losing every head-to-head.
Another put it more precisely: you can be one of five sources cited in a comparison answer and still lose, because the assistant led with someone else.
If your score cannot separate those two states, it cannot tell you whether you have a visibility problem or a preference problem. Those need completely different fixes.
Failure two: one run is not a measurement
Assistant answers are non-deterministic. The same prompt, the same model, three runs, three different brand sets. This is not an edge case; it is how the systems work.
So a tracker that fires each prompt once and plots the result over time is not drawing a trend line. It is drawing noise and labelling it a trend. As one practitioner put it, without variance shown, the chart is a liability — you will "explain" a four-point move that was never a move at all.
The related overclaim is the rank position. There is no ranked index inside ChatGPT to hold a position in. Nothing is at #3. Any tool reporting an exact AI rank is describing a structure that does not exist — and the community's assessment of who sells that is not generous. The honest version of the same question is appearance rate across a distribution of responses, reported per engine, with the spread visible.
Failure three: the score tells you nothing about what to do
Say the score is trustworthy. It is 62, you know how it was built, you know the spread. Now what?
This is where the category runs out of road. A number tells you that you are invisible. It does not tell you which page to write, which phrasing to fix, or which third-party source to go and exist in. One practitioner's summary is the fairest description of the problem we have read: diagnosis is the cheap part.
And the fix is usually not even on your own site. Someone running audits professionally reported that 93–95% of citations came from third-party sources, not the brand's own domain. Engines also do not agree with each other about where to look — ChatGPT leaning on official and reference pages, Perplexity on Reddit and YouTube, Gemini tracking Google rank more closely. "Improve AI visibility" is not one project. It is a different project per engine.
A blended cross-engine number averages that distinction away, which is exactly the distinction you needed.
The three numbers we report instead
We are not arguing that measurement is impossible. We are arguing that one number is the wrong container. We report three, and we keep them apart:
Mention rate. How often the assistant names you at all, including passing references and comparison lists. This is the number most tools report on its own.
Recommendation rate. How often the assistant puts you forward as the answer rather than listing you among others. When this diverges from mention rate, that gap is your actual problem, and it is invisible in a blend.
Agent-driven revenue. Completed transactions attributable to an agent surface, measured at the checkout rather than inferred from referral traffic. This is the number the first two exist to move.
That third one is where we part company with the category, and it is deliberate. Most AI-influenced visits arrive with the referrer stripped and land in analytics as Direct, which is why measuring at the checkout matters more than measuring traffic — the topic we cover in Agentic Commerce Traffic)">Dark Agentic Commerce Traffic. A visibility tool structurally cannot report a purchase. That is not a feature gap; it is the difference between watching a channel and running one.
What to ask any vendor, including us
Five questions. They are all fair, and a vendor who cannot answer them is selling a number, not a measurement.
- What is the prompt set, and can I see it? If it is auto-generated from keywords and you cannot inspect it, you do not know what is being measured.
- How many runs per prompt, and what is the spread? One run per prompt is noise. Ask for the variance, not just the mean.
- Are mentions and recommendations counted separately? If they are blended, the most decision-relevant signal has been averaged away.
- Is it broken out per engine? A single cross-engine percentage hides which surface you are losing.
- Which sources were cited? Without this you get a diagnosis and no treatment — and most of the treatment is off your own domain.
We publish our answers to all five on our methodology page, and we hold our own Visibility Score to them: never a bare number, always its five inputs, its run count and its per-engine spread. If anything we publish is not reproducible from what is written there, that is our error and we want to hear about it.
Because the useful version of this question was never "what is my score". It was "what do I change on Monday, and how will I know it worked". When we ran that loop with House of Zelena, the thing that moved was not a dashboard reading — it was which conversations the brand showed up inside, and what happened next.
A score is a starting point for that conversation. It has just been getting sold as the end of it.
FAQ
Continue reading
July 24, 2026
GEO vs AEO vs SEO: What Actually Changed and What Didn't
Three acronyms, three different jobs. SEO wins the crawl, AEO wins the answer box, GEO wins the citation inside a generated response. Here's what each actually optimizes for.
July 23, 2026
Why Founders Are DIY-Hacking Their AI Citation Tracking (And Where It Breaks)
Founders are already writing scripts to check if ChatGPT and Perplexity mention their brand. The instinct is right. The DIY version breaks down exactly where it starts to matter.
September 10, 2026
Zomato MCP and Swiggy MCP: India's Apps Inside AI Assistants
Zomato and Swiggy both run MCP servers, so an AI assistant can search restaurants, build a cart and place a real order without the app ever opening. Zomato shipped an official open-source server in September 2025; Swiggy followed in January 2026 across food delivery, Instamart and Dineout. Here is how each works, what Instamart and BigBasket actually did, and what the pattern tells every other brand.