AI visibility is measured with a fixed panel of prompts, run on a schedule, scored on a three-level rubric — not with a rank tracker, because generated answers have no stable positions to track. The method is less precise than rank tracking and considerably more useful than the alternative, which is inferring AI performance from organic traffic and hoping.
This is the complete methodology, including the controls that make non-deterministic output trendable and the metrics that will mislead you.
Why rank tracking does not transfer
Three properties of generated answers break the model rank tracking depends on.
There is no position. An answer either cites you or it does not. There is no third result to climb toward.
Output is non-deterministic. The same prompt run twice can produce different sources. A single observation is not a measurement.
Context shifts results. Account history, location, and session state all influence what an assistant returns, so two people running the same prompt legitimately see different answers.
What survives is frequency. You cannot know your position, but you can know how often you appear across a controlled, repeated sample — and frequency trends are actionable.
Build the panel
Write 20 to 50 prompts a real buyer would actually type. Fewer than 20 and single-prompt noise dominates; beyond 50 the manual cost outweighs the added resolution.
Cover three intents deliberately, because they fail for different reasons and the fix differs:
| Intent | Example | What failure indicates |
|---|---|---|
| Definitional | "what is answer engine optimization" | No quotable definition on your site |
| Comparative | "AEO vs SEO" | No structured comparison content |
| Commercial | "who does AEO for ecommerce" | Weak entity grounding |
Source the phrasing from real inputs — sales calls, support tickets, the questions prospects actually ask — rather than keyword tools. Prompts are spoken language, and keyword-derived phrasing does not match how people write to an assistant.
Then freeze the list. A panel that changes between cycles measures nothing. Add new prompts to a separate second panel if you need to expand coverage.
Score on three levels
Binary scoring throws away the most diagnostically useful state — being known but not quoted.
| Score | Outcome | What it tells you |
|---|---|---|
| 0 | Absent | Not retrieved for this intent at all |
| 1 | Mentioned, not linked | Entity is recognised; no page was quotable |
| 2 | Cited with a link | Full retrieval path cleared |
The distinction between 0 and 1 is the one that directs your work. A panel scoring mostly 0s is an entity grounding problem — the model does not know who you are. A panel scoring mostly 1s is an extraction problem — it knows you and cannot quote you. Those are different quarters of work, and binary scoring cannot tell them apart.
Record one more field on every score of 2: which URL was cited. That column is the most valuable data the exercise produces, because it shows you which page structures are working so you can replicate them deliberately.
The controls that make it valid
Four disciplines, all of which are the difference between a signal and a story.
- Three runs per prompt, averaged. Non-determinism means one run is noise. Three is the practical minimum for a number you can act on.
- Logged out, consistent location. Personalisation contaminates results. Use a clean session from the same region every cycle.
- Same day of month, same engines. Fix the schedule and the engine list. Adding an engine mid-series makes the trend uninterpretable.
- One person or one script. Judgement calls about what counts as a mention should be made consistently.
Monthly is the right cadence for most businesses. Weekly measures noise; quarterly is too slow to connect a change to its cause.
What to report
Two numbers carry most of the value. Citation rate is the share of panel prompts scoring 2. Mean score per intent bucket captures partial progress that citation rate alone hides — moving a bucket from 0.2 to 0.9 is real improvement even if the citation rate barely moves.
Report both against the previous cycle, with the cited-URL list attached. Resist the urge to add composite "visibility scores" that blend the two; they obscure which of the two problems you are actually solving.
Three metrics that will mislead you
Referral traffic from AI platforms. Useful but not a visibility measure. Most AI citations do not produce a click, so referral volume undercounts influence severely.
Brand search volume. Lags by months and is contaminated by every other marketing activity you run.
Any single spectacular result. The screenshot of your brand in an AI Overview is a data point of one, from a non-deterministic system. It is a nice thing to see and it is not evidence.
Frequently asked questions
Can I track AI visibility with a tool?
Several vendors now offer AI visibility tracking, and they automate the panel-and-score method described here. They are worth evaluating for scale, but understand the methodology first — the output is only as good as the prompt panel behind it, and that panel should reflect your buyers rather than a generic list.
How many prompts do I need?
Twenty to fifty. Fewer and one prompt's noise dominates the average; more and manual running becomes impractical without tooling.
How often should I run the panel?
Monthly for most businesses. Structural changes need to be recrawled and re-embedded before they affect retrieval, so weekly runs mostly measure variance rather than progress.
Why does the same prompt give different answers?
Generated responses are non-deterministic, and retrieval can surface different sources between runs. This is why the methodology requires multiple runs averaged rather than single observations.
What is a good citation rate?
There is no meaningful benchmark, because it depends entirely on your category's competitiveness and how your buyers phrase questions. The number that matters is your own, moving in the right direction across a frozen panel.
Where to start
Write ten prompts this week from your last ten sales conversations, run them once across two engines, and score them. That first run is not a measurement — it is a baseline sanity check, and it usually reveals immediately whether you have an entity problem or an extraction problem.
Then build the panel out properly and freeze it. For help constructing a panel that reflects a real buying process, that is a strategy session.
Insights from the Field
Practical guidance on SEO, AEO, and scalable growth — based on real systems, not theory.

