How We Measure AI Visibility
Our methodology in full: which engines we query, what counts as a citation, how grounding changes the meaning of a result, and the specific places we return 'we cannot tell you' rather than a plausible number.
A visibility score is a number produced by a chain of decisions — which engines, how many samples, what counts as a mention, how failures are handled. None of those decisions is visible from a dashboard, and two products can render the same number having measured very different things.
So here is ours, in enough detail to argue with. If you are evaluating tools in this category, how to audit an AI visibility vendor’s numbers is the list of questions worth asking anyone; this is us answering it about ourselves.
What we query
Seven surfaces: ChatGPT, Perplexity, Gemini, Claude, Grok, Meta AI, and Google AI Overviews.
The seventh is structurally different from the other six. It reads a search-results provider rather than a chat model, because AI Overviews is a Google Search surface rather than an assistant. When no SERP provider is configured it is excluded from scans entirely rather than silently returning empty results — a missing engine should be visibly missing, not quietly counted as an absence.
The other six are queried through model APIs. This is worth being precise about: an API call is not identical to the consumer app. System prompts, default grounding behaviour and model versions can differ from what your customer sees in the ChatGPT interface. We think API access is the right trade — it is reproducible, it is versioned, and it does not depend on scraping a UI that changes weekly — but it is a trade, and the gap is real.
Model IDs are configurable per engine but are not stored on the scan. This is a gap rather than a policy, and it is ours: the resolved model id is used to make the call and then discarded, so nothing on a scan, engine-result or answer row records which model produced a given answer. Gemini’s has been repointed twice since launch. We already decline to claim historical comparability across a model change — a new model is a new instrument — but with no id on the row we cannot even tell you where the discontinuity falls, which is the more useful half of that statement and the one we owe you.
Grounded and ungrounded, and why the distinction is everything
An engine can answer two ways, and the difference determines what a result means.
Ungrounded — the model answers from what it learned during training. This reflects the web as it was at the knowledge cutoff. It responds slowly to anything you publish, and it is the closest thing to “what does this system believe about you by default.”
Grounded — the engine performs a live web search and composes from what it retrieves. Google documents this for Gemini in its search grounding documentation; Anthropic documents web search as a provider-executed tool in its web search tool reference. Grounded results respond to new content within days.
Both are legitimate measurements of different things. Blending them without disclosure makes results uninterpretable — you publish a page, see no movement, and cannot tell whether the tactic failed or whether the engine was never going to look.
So grounding ships off. It is one global switch on our side rather than a per-engine or per-account setting, and when it is on, a grounded prompt bills a multiple of an ungrounded one — but only on the engines where search is an opt-in mechanism we chose to turn on. Perplexity and Google AI Overviews search on every call whether the switch is on or not, so their search is already inside their base rate and surcharging them would bill twice for one search. Read that consequence carefully, because it is the opposite of the intuition: with the switch off, the only grounded answers in your scan are exactly the two that are never surcharged, so there is no billing signal telling the two modes apart.
Nor is the mode written down. No scan, engine-result or answer row carries a grounding column — the closest record is the citation shape described in the next section, and it is a proxy rather than a fact, because a grounded call that retrieved nothing looks exactly like an ungrounded one.
One honest wrinkle. Anthropic’s web search is a provider-executed server tool that our AI gateway cannot pass through, so grounded Claude queries go through a direct provider client. Without that credential configured, a grounded Claude query degrades to an ungrounded call — and bills at the standard rate rather than the grounded one, because you did not receive a grounded result.
What we count as a citation
Two different things get called citations across this industry, and we distinguish them because they are not the same evidence.
A real URL. The engine performed a search, fetched a document, and told us which one. This is a source the engine demonstrably used.
An extracted domain. The engine answered from weights and named a site in its prose. Our analyzer extracted the domain. This is a domain the model produced from memory. It may not correspond to a page that exists.
Grounded scans produce the first. Ungrounded scans produce the second, and the URL field is null. We keep them apart rather than reporting a blended count.
The consequence shows up in a specific derivation. The page-level version of our Search↔AI gap analysis compares the pages Google ranks against the pages engines cite — and that requires real URLs. Run it over ungrounded scans and every ranking page would appear uncited, because there are no citation URLs to match against. That would be a confident, wrong, and extremely plausible-looking finding.
So the derivation checks whether URLs are available and, if not, reports nothing. Not zero, not “no gaps found” — it says it cannot run. A tool that returns a number in this situation is not measuring more than we are; it is measuring the absence of data.
What goes into the score
Three inputs, per engine, over the answers that engine actually delivered: mention rate (50%), position when mentioned (35%) and sentiment (15%). Each engine gets a 0–100 figure, and the Unified Score is the average across the engines that returned usable answers, weighted by how many answers each one delivered — so an engine that managed two of fifty prompts before timing out contributes two prompts’ worth of evidence, not an equal share.
Two things are deliberately outside that formula, and the omissions are as much a methodology choice as the weights. Citations carry no weight in the score. We collect them, we report them on their own surface, and we do not fold them in — a brand cited once by a low-authority blog would otherwise outscore one recommended by name and linked nowhere. Nor does your position relative to a competitor. The score measures your presence in a set of answers; making it relative would mean your number moved when a rival published something, which is a fact about them.
The load-bearing detail is what happens to failures. An engine call that errored is excluded — from the score, from the quota, and from your bill. It is not counted as an absence.
This matters more than it sounds. If a timeout counted as “not mentioned,” your score would drop whenever our infrastructure or a provider had a bad day, and you would interpret it as a visibility problem and go looking for a cause that does not exist. The same number drives billing: you are not charged for an engine that failed to answer. One decision, applied consistently on both sides.
What we refuse to show
Three things we could display and do not.
Prompt volume. We cannot source it honestly. Nobody outside the engine vendors has query logs, so every figure in the market is a panel, a search-volume proxy, or a model-generated estimate. Showing one would mean presenting an estimate with the visual authority of a measurement.
Revenue attribution. We cannot connect an AI mention to a purchase. Nobody can — the causal chain routinely runs through a direct visit weeks later with no identifier surviving. We show presence, position and share, which are observed, and we leave the commercial inference to you rather than modelling it and calling it a result.
Confidence we do not have. Every score is a sample from a non-deterministic system. Small week-over-week movements are frequently noise. We publish the arithmetic in how many runs before you trust an AI visibility number rather than letting a two-point change read as progress.
Two design decisions worth explaining
Competitor scans notify nobody. We can scan a competitor’s brand the same way we scan yours — that is how competitor benchmarking works. But a competitor scan produces a competitor’s score, and emailing you “Acme: 62/100” as an alert about your brand is a category error that reads as perfectly plausible for months before anyone notices the number was never theirs. So notifications check the scan’s subject and return nothing for anything that is not your own brand.
Sample data is labelled and can never reach billing. Before your first completed scan, the dashboard renders illustrative charts so the layout is not empty. These are labelled as sample data in the interface, are never written to the database, and cannot reach usage or credit accounting — which reads only completed scan rows. An unlabelled placeholder that a user mistakes for their own numbers is the worst possible first impression, and the second-worst is one that shows up on an invoice.
Where our numbers are weakest
The section a methodology document is actually for.
Prompt selection dominates everything. Your visibility percentage is computed over prompts someone chose. Choose flattering prompts and the number flatters you. This is the single largest source of variation between two honest measurements of the same brand, and no tool can fix it — it is a judgement call. Choosing tracked prompts is our attempt to make it a considered one.
Region is a setting, not a survey. A scan is conditioned on one region. A brand selling into eight markets and scanning one has a number about one market.
Position classification interprets prose. Deciding whether framing is favourable is a judgement a model makes, and it can be wrong on ambiguous answers. This is why we keep the underlying answers readable rather than discarding them after scoring — a classification you cannot check against the source text is not one you should defend in a meeting.
Sentiment is coarser than it looks. A three-level read on a paragraph of nuanced prose loses a lot. It is useful for detecting a change and poor for characterising a single answer.
Why publish this
Partly because we ask other vendors to. Mostly because a number whose provenance you cannot inspect is not much use: you will eventually make a decision on it, and if you cannot reason about what it measured, you cannot tell whether the decision was sound.
We would rather be argued with about a documented methodology than trusted about an undocumented one. If something here looks wrong, that is a conversation worth having — and what AEO cannot do is the broader list of limits this methodology operates inside.
Written by
Team @ LLM MetrixWe research and write about AI brand visibility, GEO, AEO, and the evolving AI search landscape.
More in Company
See how your brand appears in AI search
Track your visibility score across ChatGPT, Claude, Gemini, Perplexity, and more — free to start.
