How We Measure AI Visibility
Our methodology in full: which engines we query, what counts as a citation, how grounding changes the meaning of a result, and the specific places we return 'we cannot tell you' rather than a plausible number.
A visibility score is a number produced by a chain of decisions — which engines, how many samples, what counts as a mention, how failures are handled. None of those decisions is visible from a dashboard, and two products can render the same number having measured very different things.
So here is ours, in enough detail to argue with. If you are evaluating tools in this category, how to audit an AI visibility vendor’s numbers is the list of questions worth asking anyone; this is us answering it about ourselves.
What we query
Seven surfaces: ChatGPT, Perplexity, Gemini, Claude, Grok, Meta AI, and Google AI Overviews.
The seventh is structurally different from the other six. It reads a search-results provider rather than a chat model, because AI Overviews is a Google Search surface rather than an assistant. When no SERP provider is configured it is excluded from scans entirely rather than silently returning empty results — a missing engine should be visibly missing, not quietly counted as an absence.
The other six are queried through model APIs. This is worth being precise about: an API call is not identical to the consumer app. System prompts, default grounding behaviour and model versions can differ from what your customer sees in the ChatGPT interface. We think API access is the right trade — it is reproducible, it is versioned, and it does not depend on scraping a UI that changes weekly — but it is a trade, and the gap is real.
Model IDs are recorded per scan. We do not claim historical comparability across a model change, because a new model is a new instrument and pretending otherwise would corrupt the series.
Grounded and ungrounded, and why the distinction is everything
An engine can answer two ways, and the difference determines what a result means.
Ungrounded — the model answers from what it learned during training. This reflects the web as it was at the knowledge cutoff. It responds slowly to anything you publish, and it is the closest thing to “what does this system believe about you by default.”
Grounded — the engine performs a live web search and composes from what it retrieves. Google documents this for Gemini in its search grounding documentation; Anthropic documents web search as a provider-executed tool in its web search tool reference. Grounded results respond to new content within days.
Both are legitimate measurements of different things. Blending them without disclosure makes results uninterpretable — you publish a page, see no movement, and cannot tell whether the tactic failed or whether the engine was never going to look.
So grounding is off by default and enabled per engine, and a grounded query costs more credits than an ungrounded one. That last detail is deliberate: the distinction shows up in your billing as well as your data, which makes it hard to forget which mode produced a number.
One honest wrinkle. Anthropic’s web search is a provider-executed server tool that our AI gateway cannot pass through, so grounded Claude queries go through a direct provider client. Without that credential configured, a grounded Claude query degrades to an ungrounded call — and bills at the standard rate rather than the grounded one, because you did not receive a grounded result.
What we count as a citation
Two different things get called citations across this industry, and we distinguish them because they are not the same evidence.
A real URL. The engine performed a search, fetched a document, and told us which one. This is a source the engine demonstrably used.
An extracted domain. The engine answered from weights and named a site in its prose. Our analyzer extracted the domain. This is a domain the model produced from memory. It may not correspond to a page that exists.
Grounded scans produce the first. Ungrounded scans produce the second, and the URL field is null. We keep them apart rather than reporting a blended count.
The consequence shows up in a specific derivation. The page-level version of our Search↔AI gap analysis compares the pages Google ranks against the pages engines cite — and that requires real URLs. Run it over ungrounded scans and every ranking page would appear uncited, because there are no citation URLs to match against. That would be a confident, wrong, and extremely plausible-looking finding.
So the derivation checks whether URLs are available and, if not, reports nothing. Not zero, not “no gaps found” — it says it cannot run. A tool that returns a number in this situation is not measuring more than we are; it is measuring the absence of data.
What goes into the score
Per engine, over the answers that engine actually delivered: whether your brand was mentioned, where in the answer, the sentiment of the framing, and which sources were cited. Aggregated across engines into a 0–100 figure.
The load-bearing detail is what happens to failures. An engine call that errored is excluded — from the score, from the quota, and from your bill. It is not counted as an absence.
This matters more than it sounds. If a timeout counted as “not mentioned,” your score would drop whenever our infrastructure or a provider had a bad day, and you would interpret it as a visibility problem and go looking for a cause that does not exist. The same number drives billing: you are not charged for an engine that failed to answer. One decision, applied consistently on both sides.
What we refuse to show
Three things we could display and do not.
Prompt volume. We cannot source it honestly. Nobody outside the engine vendors has query logs, so every figure in the market is a panel, a search-volume proxy, or a model-generated estimate. Showing one would mean presenting an estimate with the visual authority of a measurement.
Revenue attribution. We cannot connect an AI mention to a purchase. Nobody can — the causal chain routinely runs through a direct visit weeks later with no identifier surviving. We show presence, position and share, which are observed, and we leave the commercial inference to you rather than modelling it and calling it a result.
Confidence we do not have. Every score is a sample from a non-deterministic system. Small week-over-week movements are frequently noise. We publish the arithmetic in how many runs before you trust an AI visibility number rather than letting a two-point change read as progress.
Two design decisions worth explaining
Competitor scans notify nobody. We can scan a competitor’s brand the same way we scan yours — that is how competitor benchmarking works. But a competitor scan produces a competitor’s score, and emailing you “Acme: 62/100” as an alert about your brand is a category error that reads as perfectly plausible for months before anyone notices the number was never theirs. So notifications check the scan’s subject and return nothing for anything that is not your own brand.
Sample data is labelled and can never reach billing. Before your first completed scan, the dashboard renders illustrative charts so the layout is not empty. These are labelled as sample data in the interface, are never written to the database, and cannot reach usage or credit accounting — which reads only completed scan rows. An unlabelled placeholder that a user mistakes for their own numbers is the worst possible first impression, and the second-worst is one that shows up on an invoice.
Where our numbers are weakest
The section a methodology document is actually for.
Prompt selection dominates everything. Your visibility percentage is computed over prompts someone chose. Choose flattering prompts and the number flatters you. This is the single largest source of variation between two honest measurements of the same brand, and no tool can fix it — it is a judgement call. Choosing tracked prompts is our attempt to make it a considered one.
Region is a setting, not a survey. A scan is conditioned on one region. A brand selling into eight markets and scanning one has a number about one market.
Position classification interprets prose. Deciding whether framing is favourable is a judgement a model makes, and it can be wrong on ambiguous answers. This is why we keep the underlying answers readable rather than discarding them after scoring — a classification you cannot check against the source text is not one you should defend in a meeting.
Sentiment is coarser than it looks. A three-level read on a paragraph of nuanced prose loses a lot. It is useful for detecting a change and poor for characterising a single answer.
Why publish this
Partly because we ask other vendors to. Mostly because a number whose provenance you cannot inspect is not much use: you will eventually make a decision on it, and if you cannot reason about what it measured, you cannot tell whether the decision was sound.
We would rather be argued with about a documented methodology than trusted about an undocumented one. If something here looks wrong, that is a conversation worth having — and what AEO cannot do is the broader list of limits this methodology operates inside.
Written by
Team @ LLM MetrixWe research and write about AI brand visibility, GEO, AEO, and the evolving AI search landscape.
More in Company
See how your brand appears in AI search
Track your visibility score across ChatGPT, Claude, Gemini, Perplexity, and more — free to start.
