The AI Visibility Metrics That Actually Matter
Mention rate, citation rate, share of voice, share of model — AEO has its own vocabulary of metrics. Here's what each one measures, when to use it, and how they fit together.
In traditional SEO the metrics were settled: rankings, impressions, clicks, conversions. AEO is younger, and its vocabulary is still forming — which means a lot of teams either measure the wrong thing or can’t agree on what “good” looks like. This guide lays out the metrics that matter for AI visibility, what each one tells you, and how they stack into a system.
It also covers the part that gets skipped: what each metric cannot tell you. Half of measuring AI visibility well is knowing which absences are findings and which are just missing data.
Start with the denominator: your query universe
Before any metric, you need a defined set of questions to measure against — your query universe. Every metric below is “out of how many tracked queries?”, so a representative, consistently-tracked query set is the foundation. Measure against a moving or arbitrary set and your trends mean nothing, because you’ve changed the denominator and the numerator at the same time.
This is also why AEO needed new metrics at all rather than borrowing search’s. A generated answer has no ranked list to read a position off. The academic work hit the same wall: the 2023 paper GEO: Generative Engine Optimization had to define its own impression metrics — weighting a source by how prominently it features in the generated text — rather than reuse click-and-rank measures. Every metric below is a variation on that problem.
Mention rate: are you in the room?
Mention rate is the share of relevant answers in which your brand is named at all. It’s the most basic AEO metric and the right starting point: if you’re not mentioned, nothing else matters yet. Track it per engine, because presence varies a lot between ChatGPT, Perplexity, Gemini, Claude and the rest.
The failure mode is treating a low mention rate on one run as a finding. Generated answers are sampled, so mention rate has genuine run-to-run variance. Measure repeatedly and compare distributions, not single numbers.
Position: how prominent are you?
Being mentioned last in a list of eight is not the same as being the first recommendation. Position, or prominence, captures where and how emphatically you appear. Two brands with identical mention rates can have very different real-world impact. See understanding position drift for how this moves over time.
One measurement detail worth carrying over from classic search reporting: when you average a position across many queries, weight it by exposure rather than taking a plain mean. In our search performance view, average Google position is impression-weighted — SUM(position × impressions) / SUM(impressions) — because a three-impression query should not pull as hard on the headline number as a thirty-thousand-impression one. A plain mean is the number a marketer will screenshot and later have to defend.
Citation rate: are you the source?
For citation-first engines like Perplexity and Google AI Overviews, citation rate measures how often your pages are credited as the source behind an answer — not just whether your brand is named. It’s one of the most concrete AEO metrics because the citations are visible and attributable.
Here is the catch, and it’s the single most common way AI-visibility reporting goes wrong. A “citation” is not always a URL. When an engine is queried as a plain chat model with no live web search, there are no links in the response — an analyzer can extract a domain from the prose, but not the page. In LLM Metrix, an ungrounded scan stores a null URL for exactly that reason, and any URL-level analysis is suppressed rather than run on empty data.
That guard exists because the alternative is worse than useless. With zero citation URLs, a naive “which of my ranking pages are not cited?” report returns every ranking page — a confident, precise-looking finding that is entirely an artefact of the data being absent. If your tool reports page-level citation gaps, ask it whether the underlying scan was grounded. See how AI engines cite sources.
Share of voice: how do you stack up?
Absolute metrics tell you how you’re doing; competitive metrics tell you whether you’re winning. Share of voice measures your slice of all brand mentions in your category across tracked queries. These are the numbers executives care about, because they answer “are we beating our competitors in AI answers?”
The honest caveat: share of voice is relative to your query set, not to the whole of AI. A competitor’s share in your data reflects how often they appear in the questions relevant to your buyers — not their overall presence everywhere. That’s usually the more useful number, but it is not the number the phrase implies, so say which one you’re reporting.
Sentiment and accuracy: is the mention good?
A mention isn’t automatically a win. Sentiment captures whether you’re described positively, neutrally or negatively; accuracy captures whether the description is even correct. An inaccurate or negative mention can be worse than no mention, which is why brand-safety monitoring belongs in your metric set rather than in a separate quarterly review. See why does AI get my brand wrong.
A composite score has to be careful here. In our own scoring, an engine’s score blends mention rate, prominence and sentiment — but the sentiment term is gated on there being at least one mention. That sounds pedantic and isn’t: an early version let a neutral sentiment default through for brands with no mentions at all, which gave a brand that no engine ever names a non-zero floor instead of zero. The metric was quietly reporting reputation for a brand that had no presence to have a reputation about.
Coverage: the metric nobody reports
Every metric above is a ratio over answered questions. The interesting question is what happened to the ones you couldn’t answer.
Our Search↔AI gap view is built around this. It joins Google Search Console queries to scanned prompts and reports two disagreements: queries where you rank on page one but no engine mentions you, and prompts where engines mention you but Google doesn’t rank you. What it deliberately does not do is convert missing data into findings:
- A Search Console query that matches no tracked prompt is not “AI ignores you.” Nobody asked the engines about it.
- An AI-mentioned prompt matching no Search Console query is not “you don’t rank.” Google withholds low-volume queries entirely — its own deep dive on Search Console performance data explains that anonymized queries are excluded from the query table while still counting toward the chart totals, which is why summing the query rows never reaches the total.
Both categories are counted and surfaced as coverage instead. It’s a less exciting number than a finding, and it’s the one that keeps the findings trustworthy. A dashboard that hides its coverage is a dashboard that will eventually hand you a claim your customer can disprove by opening Search Console in the next tab.
Referral traffic: the click-through floor
Finally, AI referral traffic measures visitors who actually click through from an AI answer. It’s valuable owned data, but read it correctly: it captures only click-throughs and misses zero-click impact entirely, so treat it as a floor on your impact rather than a measure of it.
How they fit together
Think of these as a funnel of increasingly demanding questions:
- Are you mentioned? (mention rate)
- Are you prominent? (position)
- Are you the source? (citation rate — grounded scans only, for URLs)
- Are you winning? (share of voice, within your query set)
- Is the mention good and true? (sentiment, accuracy)
- Is it driving action? (referral traffic plus lift attribution)
- How much of this do we actually have evidence for? (coverage)
No single number captures AI visibility. Teams that measure well track a small, consistent set across these dimensions, watch the trend rather than any single snapshot, report coverage alongside findings, and tie movements back to specific actions without over-claiming causation. Start with a defined query universe and mention rate, then layer in the rest as your programme matures.
Written by
Team @ LLM MetrixWe research and write about AI brand visibility, GEO, AEO, and the evolving AI search landscape.
See how your brand appears in AI search
Track your visibility score across ChatGPT, Claude, Gemini, Perplexity, and more — free to start.
