01 Metrics
Measuring AI Visibility Across Platforms
How to measure brand presence in AI answers: a fixed prompt set, repeated runs across several models, and the five signals worth recording.
- October 8, 2025
- 3 min
Measuring AI visibility means running a fixed set of prompts against several AI answer engines, repeating each one, and recording which brands the answers name and which sources they cite. The GoAnswers diagnostic runs 30-40 prompts across 4 models with 3 iterations each, which is what turns a screenshot of one answer into a rate you can compare next month.
One answer is not a measurement
Answer engines are not deterministic. Ask the same question twice and the brands named can differ, because the model samples its own output and, on the engines that search live, retrieves a different set of pages.
Running each prompt three times per model is the minimum that separates "we appear" from "we appeared once". A brand named in two runs out of three has a frequency. A brand named in one screenshot has an anecdote.
What gets recorded
Five signals come out of the same set of runs:
- Share of Recommendation — how often the engine recommends your brand relative to competitors across the prompt set.
- Citation Share — the proportion of answers that reference your pages as a source.
- Top3 — whether your brand is among the first three named.
- Co-mentions — which competitors appear alongside you inside the same answer.
- Sentiment — how the answer characterises you when it does name you.
Citation and recommendation are different outcomes and they move independently. A model can cite your comparison page as evidence and recommend a competitor in the same breath.
Write the scoring rules down before the first run. Decide what counts as a recommendation rather than a passing mention, whether a brand named only inside a quoted source counts, and how you treat an answer that lists ten options. Two people scoring the same transcript should arrive at the same number.
Choosing the prompts
The set should be the questions buyers actually type, not the keywords you would bid on. In practice that means category prompts ("best [category] for [use case]"), head-to-head comparisons, "alternatives to [competitor]", and problem-shaped questions that never mention a product.
Include the prompts you expect to lose. A set built only from questions you already win returns a high number and no information.
Then freeze it. If the prompt set changes between cycles, the movement in the numbers is the movement in the prompts.
Fix the exact wording too, not just the topic. Paraphrasing a prompt changes what gets retrieved, so a reworded prompt is a new prompt and starts its own baseline.
Why four models
Each engine weights different signals, so a single-model number does not generalise.
- ChatGPT combines training data with live search; brand mentions and authority carry weight, and it crawls as GPTBot and OAI-SearchBot.
- Perplexity is the citation engine: fresh sources with linkable structure, fetched by PerplexityBot.
- Google AI Overviews draws on Google's index, so conventional SEO, passage-level structure and schema still apply, with the direct answer in the first paragraph.
- Claude leans on trusted sources and entity consistency, crawling as ClaudeBot and Claude-SearchBot.
Gemini overlaps with AI Overviews inside Google's ecosystem, and Copilot sits on the Bing index. A brand can hold Top3 in Perplexity and be absent from AI Overviews. That gap is the finding, and it points at different work in each case.
Keep the answers, not just the scores
Save the full answer text, the URLs cited, the date and the model used. The scores tell you where you stand; the transcripts tell you why.
Reading them is where the specifics come from: the competitor page cited on every category prompt, the outdated claim about your pricing that three models repeat, the answer that names you but files you under the wrong category. None of that survives compression into a percentage.
Cadence, and what this cannot tell you
Re-run monthly and compare quarterly, holding the prompt set, the model list and the iteration count constant while you do it.
Two limits are worth stating plainly. This is a sample, not a census: it describes the prompts you chose, on the dates you ran them. And no engine publishes its ranking factors, while models are updated without notice, so a drop can come from a model change rather than from anything you did. The dated transcripts are what let you tell those apart.
It also measures visibility, not revenue. Share of Recommendation is a leading indicator, and it has to be read next to whatever your analytics show for direct and assisted traffic.