How we measure AI visibility
How we measure, how certain the figures are, and where our statement stops.
Last updated: 15. Juli 2026
What is measured
We measure two things that deliberately stay apart and are never combined into a single figure.
- Visibility
- Does a business come up when someone asks a question that does not contain its name? Something like „good coffee shop in Manchester“. That is the number that matters when it comes to new customers.
- Brand knowledge
- What does a model say about a business when you ask about it directly? Here the name is in the question, so the model is bound to mention it. A brand can do excellently here and still be invisible in its category - a combined figure would hide exactly that.
How often it is measured
A single measurement cannot be evaluated. The same question produces different answers across two runs - a true mention rate of 50 per cent can appear as 0 or as 100 per cent.
We therefore measure repeatedly and evaluate over a rolling 14-day window. One run on fourteen days is as accurate as seven runs on a single day, costs the same, and additionally shows how visibility moves over time.
Channels with a lower measurement density are marked as such in the tool. Their figures are less certain, and it says so there.
Which AI systems are queried
ChatGPT, Google AI Overviews, Gemini, Claude, Perplexity and the AI answer box in Bing.
For clarity: with Bing we measure the AI box in the search results, not the Copilot chat product. Not every provider makes that distinction, but it matters - these are different systems giving different answers.
For brand knowledge we query without web search. What is measured is then what the model holds, not what it currently finds online.
Why every statement carries evidence
Every attribute we display is backed by a verbatim quote from the AI answer. The quote is checked against the original text. Anything that cannot be evidenced is discarded and never appears in the first place.
The summarising sentences, too, are derived solely from the figures already calculated, not from the raw answers. Running a language model over AI answers would layer a second source of error on top of the first.
How we show uncertainty
- Always with a denominator
- „3 of 18 answers“ rather than „17 per cent“. At eighteen measurements the actual rate sits within a range of roughly twenty percentage points - a round percentage conceals that.
- Confidence interval
- Every rate comes with a range within which the true value lies (Wilson interval). It is shown alongside.
- Data quality indicator
- Green from fourteen measurements, amber from seven, below that grey with a note that the basis is not yet sufficient for a statement.
- No trend drawn from noise
- Where the confidence intervals of two periods overlap, we show no arrow but state that no change can be demonstrated.
What these numbers are not
The most important section on this page. We deliberately draw the line where our measurement stops.
- No extrapolation to visitors or revenue
- There is no evidenced link between an AI mention and actual usage. Figures such as „this brings you X visitors“ would be guessed, not measured.
- No personalised view
- AI answers depend on the history and context of the person asking. What an individual user sees is beyond our measurement - we measure a neutral query.
- No position from a single measurement
- Rank within an answer is the least stable of all quantities. We report it as a median with a range, never as a single value and never as a mean.
- No complete coverage
- We measure a selection of questions, not everything that is ever asked about a business. The selection is visible and editable in the tool.
- No guarantee of improvement
- We show where a business stands and what stands out. Whether and when a model changes its answer is outside our control.
When the AI states something false
AI systems demonstrably invent details more often about businesses that little has been written about publicly - which means precisely about small and medium-sized operations. Wrong locations, wrong services, confusion with other companies.
We collect verifiable factual claims separately from the attributes and make them flaggable. That makes visible what a business ought to correct.