You have decided you need a tool. Every vendor’s homepage shows the same screenshot: a line going up, a list of engines, a share-of-voice donut. None of it tells you whether the numbers underneath are measurements or estimates.
This is a framework, not a shortlist. Run it against any vendor, including us. Each question has a specific technical answer a salesperson can give in one sentence, and the ones who cannot are telling you something. If you are still deciding what categories of tooling you need, start with how to choose AEO tools and build your stack.
1. Engine coverage — and whether the queries are grounded
The engine count on the pricing page is the least interesting number here. The real question is how each engine is queried.
There are two ways to ask ChatGPT what the best CRM is: send the prompt to a plain chat model answering from training data, or to a search-enabled model that retrieves live pages and returns its sources. Either is legitimate — but a tool doing the first and calling it “monitoring what ChatGPT tells your buyers” measures a model’s memory, not the product your buyers use.
Ask: for each engine, is the query grounded in live web search, or a plain model call? The honest answer is per engine and conditional. In our own system, Perplexity’s Sonar searches on every call regardless of settings; Google AI Overviews is a read of the actual SERP surface, not a model call; Meta AI has no web search capability and is always ungrounded; ChatGPT, Gemini, Grok and Claude ground only when grounding is enabled. Four different answers inside one product. A vendor who says “we query all engines live” across a nine-engine list is either doing something hard or not distinguishing the cases.
Grounding costs real money, which is why it is rarely on by default anywhere: Anthropic prices its web search server tool at $10 per 1,000 searches on top of token costs (Anthropic web search tool docs). If a vendor grounds every engine on every run at a flat monthly price, ask how. See multi-engine monitoring and our multi-engine feature page.
2. Citations: real URLs, or domains scraped out of prose?
This follows from the first question and almost nobody asks it.
A grounded engine returns structured sources — Anthropic’s tool response carries a url, title and cited text for each. An ungrounded engine returns no sources, so a platform that still shows a “citations” table built it by having a second model read the answer and pull out domains mentioned in prose. That is a reasonable fallback, and we use it. It is not the same data: a domain named in a sentence is not a link the engine retrieved.
The tell is at the URL level. In our schema an ungrounded citation row stores a domain and a null URL, because there is no URL to store. Ask: can I see the full cited URL, or only the domain? If the answer is “domain,” the citations are extracted, whatever the label says — and page-level analysis of which of your pages engines cite is impossible. See how AI engines cite sources and citation intelligence.
3. Determinism and sampling
Run the same prompt through the same engine twice and you get two different answers. That is how the systems work, not a bug — which means every number a visibility platform shows you is a sample, and the sampling design is a methodology choice the vendor made for you.
Ask: how many times is each prompt run per period? Is a reading one run or an aggregate? Do you show variance or only a point value? If mention rate moved from 40% to 55%, is that real or inside the noise band?
Be suspicious of a tool that never shows uncertainty: sampling once a day across ten prompts and reporting a two-decimal score presents noise as precision. Why queries return different results explains the mechanics.
4. Data portability
Assume you will leave. What comes with you?
- An API. Read access to projects and scans over a documented, authenticated endpoint. Ours is a bearer-token REST API at
/api/v1/projectsand/api/v1/scans, throttled to 120 requests per minute per key. - Flat exports. CSV of the underlying rows, not a PDF of a chart.
- A full account export. GDPR gives you a right to your personal data; ask whether it is self-serve or a support ticket.
The failure mode is a tool whose only export is a screenshot-quality report. History you cannot extract is not an asset.
5. Retention, and whether it differs by plan
Retention is where the marketing number and the deletion job usually disagree. Ask for the window per plan, and what happens when you downgrade.
Ours: 7 days on free, 90 days on team, 1 year on business, 2 years on agency — published on the pricing page and enforced by a scheduled job reading the same constant the pricing table renders. Seven days on free is short: long enough to evaluate the product, not to build a trend line.
6. Does it join to anything you already trust?
AI visibility data is new, unfamiliar and unauditable by your CFO. The strongest thing a platform can do is connect it to a source you already believe — usually Google Search Console: first-party measurement of a property you verified, with history backfilled on connect. Ask a vendor how far back that backfill actually reaches, and for which rows — the answer is usually two windows, not one. Daily property totals typically come back much further than per-query rows, because query rows are bulkier and their long tail is withheld regardless. It matters because any query-level join a tool builds on top is bounded by the shorter window, not the headline number.
Joining it to AI answer data produces the two findings neither dataset yields alone: queries where you rank on page one but no engine mentions you, and prompts where engines mention you but you do not rank. Ask whether a vendor performs that join or merely shows both datasets side by side. Google’s guidance on AI features and your website sets out what Search Console does and does not report about AI surfaces.
7. What does it refuse to claim?
This best predicts whether you will still trust the tool in six months.
Search Console omits queries made by very few people to protect user privacy; those anonymized queries never appear in the query table (Search Console performance report). So when a platform matches Search Console queries against AI prompts, a query with no matching prompt does not mean “AI ignores you” — nobody asked. A prompt with no matching query does not mean “you do not rank” — Google may have withheld the term.
A trustworthy tool counts those cases as coverage. A weak one reports them as findings, because a finding demos better than a caveat. Three questions expose the difference:
- With no data for a prompt, does the dashboard show a zero, or a gap?
- If citation URLs are unavailable (criterion 2), does the page-level report degrade to nothing, or show every page as “not cited”?
- What is your false-positive story when our brand name is also a common word?
We answer the second by returning nothing: without real URLs, a page-citation comparison would mark every ranking page uncited — confident, wrong, and very plausible-looking.
8. Team and access model
Easy to defer until it bites. Ask what seats cost above the included allowance; whether a read-only role exists for stakeholders who should not trigger paid scans; and whether multiple brands or clients get genuinely separate workspaces with separate billing, or one account with a dropdown. Agencies should test this hardest — see our agency solution page for the shape we chose (owner/admin/member/viewer roles, workspace-scoped billing).
Where you should choose something else
A guide concluding “and therefore buy us” is an ad. Specific cases where we are the wrong tool:
- You need Microsoft Copilot, Amazon Rufus, Apple Intelligence, DeepSeek or You.com. We track seven surfaces: ChatGPT, Perplexity, Gemini, Claude, Grok, Meta AI and Google AI Overviews. Our library covers the others because they matter to some brands, but the product does not query them. If Copilot inside Microsoft 365 is where your buyers are, buy accordingly.
- You want AI Overviews with zero setup. Our Google AI Overviews engine reads a SERP API and needs a configured SERP provider. Without one it is excluded from scans rather than silently faked.
- You want everything grounded, always, at a flat price. Grounding is off by default here, and grounded queries on opt-in search engines bill at a multiple of the base credit rate (2× by default). Price it out first.
- You want a long free-tier runway. Free is one domain, one seat, 10 tracked prompts refreshed once a month, and 7-day retention. It runs forever and needs no card — but it is a way to see the product on your own brand, not a way to monitor it.
Sanity-check any vendor before you sign
You do not need an account — ours or anyone’s — to test a claim. Run your brand through the free tools and compare the output to the vendor’s demo:
- AI visibility score — a baseline to hold a vendor’s number against.
- Citation checker — see what “citations” look like for your domain.
- Hallucination checker — check whether engines get your facts right before paying anyone to tell you.
- Competitor compare — test a share of voice claim against a second source.
Then do the thing that matters most: pick five prompts you care about, ask each vendor to show you those five, and open the engines yourself in another tab. A manual AI visibility audit takes an afternoon and is the only benchmark genuinely yours.
Frequently Asked Questions
What is the single most important question to ask an AI visibility vendor?
Whether each engine is queried with live web search or as a plain chat model. Grounded and ungrounded queries produce different answers and different citation data, and the distinction is usually per engine rather than product-wide. A vendor who cannot answer engine by engine is probably not measuring what you think.
How can I tell if a tool’s citations are real?
Ask whether you can see the full cited URL or only the domain. Grounded responses return structured sources with URLs and titles; ungrounded ones return nothing, so any citation table built on them comes from a model extracting domain names out of prose. Domain-only citations cannot support page-level analysis.
Why does the same prompt give different numbers on different days?
AI engines are non-deterministic, so every visibility metric is a sample. Ask how many runs sit behind each reading and whether variance is shown — a small sample reported to two decimals presents noise as precision.
Should I care about retention if I am only running a pilot?
Yes. Retention differs by plan and free tiers are often short. A 7-day window is enough to see a product work but not to build a trend line, so a three-month free-tier pilot loses most of its history. Confirm the window per plan, and what happens if you downgrade.
