How to Audit an AI Visibility Vendor's Numbers
Every tool in this category will show you a confident dashboard. Nine questions separate the ones measuring something from the ones generating something — and we've answered all nine about ourselves at the end.
AI visibility tooling is a young category, which means two things. The methodologies vary enormously, and almost none of them are visible from the dashboard.
Two products can show you “Visibility Score: 64” where one queried seven engines a hundred times and classified every answer, and the other asked a language model to estimate how visible your brand probably is. Both render the same number in the same font. The buyer cannot tell them apart by looking.
These are the questions that tell them apart. They are worth asking of any vendor, including us — and because it would be poor form to publish a test we had not sat, our own answers are at the bottom.
1. Are you querying real engines, or simulating them?
The foundational question. Some tools query the actual assistants. Others ask one language model to predict what other engines would say, which is cheap, fast, and measuring nothing about the world.
There is a middle case worth catching: querying a model via an API is not always the same as the consumer product. The API version may have different grounding behaviour, a different system prompt, and different defaults than what your customer sees in the app. That gap is legitimate and should be disclosed, not papered over.
Follow-up: which specific engines, and can I see a raw answer exactly as returned?
2. Is the answer grounded or ungrounded?
Whether an engine performed a live web search before answering changes what the result means, fundamentally.
An ungrounded answer reflects what the model absorbed during training — slow to change, reflecting the web as it was at the knowledge cutoff. A grounded answer reflects live retrieval — responsive to your content within days. The distinction is documented by the vendors themselves: Google describes search grounding in its Gemini API documentation, and Anthropic documents web search as a server-side tool in its web search tool reference.
If you publish a page and the tool reports no change, you need to know which mode was used before drawing any conclusion. Ungrounded results should not respond to new content. A vendor blending both into one score without disclosure has made your results uninterpretable, and this single question separates most serious tools from most unserious ones.
3. What does “citation” mean here?
This is where the largest silent discrepancy lives.
When an engine performs a real web search, it returns actual URLs — a genuine citation with a link. When it answers from weights, there is no link, and tools recover “citations” by extracting domain names mentioned in the prose.
Both get called citations. They are not the same evidence. One is a document the engine demonstrably used; the other is a domain the model happened to name, which it may have produced from memory and which may not resolve to a real page at all.
Follow-up: what share of my citations are real URLs versus extracted domains? A vendor that cannot break this down is reporting a blended figure whose composition it does not track.
4. How many samples per data point?
Engines are non-deterministic. The same prompt run twice can return different brands. So every figure is a sample, and its reliability depends on the sample count.
Ask: how many observations underlie the headline score? How many runs per prompt per engine? Is a reported week-over-week change larger than the sampling noise?
A vendor that has not thought about this will not have an answer. One that has will tell you their sample size and be candid that small movements are not interpretable — see how many runs before you trust an AI visibility number for the arithmetic.
5. What exactly goes into the composite score?
Every tool has a proprietary 0–100 score. Reasonable — the underlying data is multidimensional and executives need one number.
The question is whether the composition is documented. Which inputs, weighted how? Does a mention count the same as a recommendation? How is a failed engine call handled — excluded, or scored as an absence?
That last one matters more than it sounds. If engine timeouts count as “not mentioned,” your score drops when the vendor’s infrastructure has a bad day. You will interpret it as a visibility problem and go looking for a cause that does not exist.
Follow-up: if I disagree with your weighting, can I see the components?
6. How do you handle regions and languages?
Answers vary by market — genuinely, not incidentally. A US-conditioned query and a German one return different brands for the same question.
Ask which region a scan represents and how it is set. A tool that silently defaults everything to a US context and reports a global figure is giving an international brand a number about one market with a label implying all of them.
7. What happens when a model updates?
Model releases can change what an engine says about an entire category overnight, in ways no action of yours caused. It is the closest analogue to an algorithm update, and it is the largest source of unexplained movement in this data.
Ask whether the vendor tracks which model version produced each result, and whether historical data is comparable across a model change. If your March and September numbers came from different models, comparing them is comparing two different instruments. Navigating AI model updates covers the practical handling.
8. Where does prompt volume come from — if you show it?
If a vendor displays how often a prompt is asked, ask for the source. There are only three honest answers: a user panel, a public proxy such as search volume, or a model-generated estimate. Each implies a specific limitation, and none of them is a query log, because nobody outside the engine vendors has one.
A crisp answer is a good sign regardless of which of the three it is. An evasive one means the third. There is no keyword volume for AI prompts covers why the data cannot exist.
9. What do you refuse to claim?
The most revealing question, and the least answerable by a script.
Every serious methodology has known limits. A vendor who has thought carefully can name theirs immediately — what their data cannot support, which conclusions they decline to draw, where the error bars are widest. A vendor who cannot has either not examined their methodology or would prefer you did not.
Watch specifically for anyone claiming to guarantee placement in AI answers, or to have you “cited within 30 days.” Nobody controls model output. Those are not aggressive claims; they are claims about a mechanism that does not exist, and can I pay to appear in AI answers covers why.
Two cheap tests that beat any demo
Ask for a raw answer. Not a score, not a chart — the literal text an engine returned for one of your prompts, with the engine and timestamp. Real measurement produces artifacts. A vendor who cannot show one is generating rather than observing.
Check a claim you can verify. Pick a prompt where you already know the answer — one where you are certain a competitor dominates, or certain you are absent. Ask the engine yourself and compare. Any tool can look plausible on data you cannot check.
Our own answers
Applying the list to LLM Metrix, since publishing it otherwise would be a bit rich.
- Real engines. Seven surfaces — ChatGPT, Perplexity, Gemini, Claude, Grok, Meta AI, and Google AI Overviews, the last reading a SERP provider rather than a chat model. Queried through model APIs, not scraped from consumer apps. Multi-engine monitoring is where the per-engine split lives, because averaging seven engines into one number hides exactly the differences worth acting on.
- Both, and the mode is recorded. Grounding is off by default and enabled per engine. A grounded query costs more credits than an ungrounded one, so the distinction is visible in your billing as well as your data.
- Both kinds, distinguished. Grounded scans record real URLs. Ungrounded scans record domains extracted from prose, and the URL field is null. Derivations that require real URLs — page-level gap analysis — report that they cannot run rather than silently treating null as absence.
- Configurable, and honestly bounded. Scan cadence is a plan feature; the sample is whatever you have accumulated. We publish the arithmetic above rather than implying more precision than the sample supports.
- Documented components. Mention rate, position, sentiment, citations, per engine. Errored answers are excluded from both the score and the credit charge — you are not billed for an engine that failed, and it does not count against you.
- Per-project region, which conditions the queries and the discovery prompts.
- Model IDs are configurable and recorded per scan. We do not claim historical comparability across a model change; that is a real discontinuity and pretending otherwise would corrupt the series.
- We do not show prompt volume. We cannot source it honestly, so we do not display it.
- Documented at length. We cannot guarantee a citation, cannot edit model weights, cannot attribute revenue to an AI mention with confidence, and cannot monitor a domain nobody signed up for. What AEO cannot do is the long version, and how we measure AI visibility is the methodology in full.
The point
This list is not really about vendor selection. It is about knowing what your own numbers mean, which matters just as much after you have bought something as before.
A tool with a modest methodology, clearly described, is more useful than an impressive one you cannot interrogate — because you can reason about the first one’s limits and you will eventually make a decision on the second one’s number without knowing what it measured. Choosing an AI visibility platform covers the feature comparison; this is the part that comes first.
Written by
Team @ LLM MetrixWe research and write about AI brand visibility, GEO, AEO, and the evolving AI search landscape.
See how your brand appears in AI search
Track your visibility score across ChatGPT, Claude, Gemini, Perplexity, and more — free to start.
