METHODOLOGY
How we measure, and what we refuse to claim.
Every number we publish carries its prompt set, its run count, its engine, and its date. Here is the method behind them - including where it stops working.
AI visibility measurement has a credibility problem, and it is deserved. Most of the category reports a single blended percentage with no methodology attached, which means two tools can both tell you “40% visibility” while measuring entirely different things. A score you cannot interrogate is not a measurement.
So this page exists before the numbers do. If anything we publish is not reproducible from what is written here, treat that as our error and tell us. For the argument behind it, read why a single AI visibility score is a vanity metric.
Three numbers, never one.
The single most common failure in this category is collapsing these into one figure. They move independently, and only the third one pays for anything.
01
Mention rate
How often does the assistant name you at all?
The share of runs in which the brand appears anywhere in the answer, including passing references and comparison lists. This is the number most tools report on its own.
02
Recommendation rate
How often does it actually tell the shopper to buy you?
The share of runs in which the brand is put forward as the answer, not merely listed. Being one of five names in a comparison is a mention, not a recommendation. These two numbers routinely diverge, which is why we never blend them.
03
Agent-driven revenue
Did anyone buy?
Completed transactions attributable to an agent surface, measured at the checkout rather than inferred from referral traffic. This is the number the other two exist to move, and the one a visibility tool structurally cannot report.
The rules we hold ourselves to.
We fix the prompt set before we measure, and we publish it
Every engagement starts with a prompt set built from real buyer questions in the category - category queries, comparison queries, and brand-specific queries. The set is agreed before the first run and held constant, because a prompt set that changes between runs makes the trend meaningless. You get the list.
We run every prompt multiple times and report the spread
Assistant answers are non-deterministic. The same prompt, on the same model, in the same hour, returns different brand sets. A tracker that fires each prompt once and draws a line is plotting noise. We run each prompt repeatedly per engine and report the distribution, not a single reading.
We report per engine, never blended
ChatGPT, Perplexity, Gemini and Claude do not draw on the same sources and do not behave alike. A single blended percentage across all of them hides the only actionable detail - which engine you are losing, and to whom. Every figure we hand over is broken out by surface.
We do not report a rank position
Nobody ranks #3 in ChatGPT. There is no ranked index to hold a position in. Any tool reporting an exact AI rank is describing something that does not exist. We report appearance rate across a response distribution, which is the honest version of the same question.
We record which sources the answer cited
For every run we log the cited sources, not just whether the brand appeared. This is the half that tells you why, and it is usually the half that points off your own domain - to a listicle, a review platform, a forum thread, or a trade publication you are absent from.
Every number carries n and a date
Prompt count, runs per prompt, engines covered, and the capture date travel with every figure we publish, on this site and in client reporting. A percentage without those four things is not a measurement.
What this method cannot tell you.
Any measurement worth trusting states its own limits. These are ours.
- We measure a prompt set, not the whole world. A prompt set is a sample of buyer intent, chosen deliberately. Change the set and the numbers change - which is why we hold ours constant and show you what is in it.
- Model updates move results, and not because you did anything. When a provider ships a model change, brand sets shift across the board. We date every capture so a step change can be told apart from a trend.
- Assistant referral traffic is largely invisible in analytics. Most AI-influenced visits arrive stripped of their referrer and land in GA4 as Direct. That is why we measure at the checkout and treat branded search as a supporting signal, not the other way round.
- Appearing is not the same as being chosen, and being chosen is not the same as being bought. We separate the three because collapsing them flatters the dashboard and costs you the deal.
- Ad-library observations are observed ads, not all ads. Where we cite advertiser data, it reflects creatives captured during a stated window - it is a lower bound, never a census.
Where our published data comes from.
Market demand figures come from Google Ads search-volume data pulled via DataForSEO, with the pull date stated wherever they appear. Our own search performance comes from Google Search Console, which we prefer over any third-party rank estimate for our own site - third-party indexes routinely miss the queries that actually send us traffic. Advertiser and creative observations come from ad corpora captured on stated dates, and are reported as lower bounds. Case-study numbers come from the brands themselves, with a named contact, a timeframe, and a baseline.
The working through-line: our research shows the numbers, case studies show what moved, and the ChatGPT Ads library shows the channel we measure most closely.