How the numbers are made
Scans, scoring, confidence bands, and the limits of what an AI visibility measurement can honestly claim.
What a scan runs
Each scan asks your tracked prompts across the engines you selected (ChatGPT, Perplexity, Gemini, Claude, Grok, Meta AI, DeepSeek; up to four on paid plans) plus Google's AI surfaces. Each engine answer is graded by a separate analyzer pass for brand mention, rank, sentiment, cited sources, and factual claims about your brand. A scan with any errored engine still bills only for delivered answers; wholly failed scans bill nothing.
Confidence bands, not vibes
A mention rate computed over n answered checks is a binomial proportion, so headline mention rates carry their 95% Wilson score interval, rendered as “24% ±13” style bands that widen when few answers back them. Small scans show wide bands on purpose: that width is the real uncertainty of the measurement, not noise to be styled away. Bands never collapse to zero at 0% or 100%, because those are exactly the scores where pretending certainty is most damaging.
When a band is wider than you want to report against, Deep Verify re-asks each tracked prompt several times across your engines and folds the repeats into the same interval. The band narrows with measured samples (roughly with the square root of how many answers back it), never by assumption: a run whose rounds errored bills and counts nothing, so width only tightens when real answers arrived. Repeats measure answer-to-answer variance for YOUR prompts on THIS day; they do not average away provider drift between weeks, which is what trends are for.
What one scan cannot tell you
Model answers vary run-to-run, by region, and as providers update. One scan-day is one sample of that distribution. Movement between two scans mixes real change with model variance. That is why trends exist (each additional scan is another sample of that distribution) and why we refuse claims a single scan cannot support: a prompt nobody asked is not “AI ignores you,” and a missing Search Console query is not “you don't rank.” Where coverage limits a conclusion, we report the coverage instead of the conclusion.
Scoring
The 0–100 visibility score blends presence (mention rate), position (ranked placements outweigh fine-print mentions, with the same tier boundaries used everywhere ranks render), citation share against tracked competitors, and sentiment. The formula is deterministic from stored answers and frozen per scan, so history never re-scores under your feet.
Data handling
Retention windows per plan (7 days to 2 years) are the same constants the pricing page renders: that is how long a scan stays readable in the dashboard. Benchmark percentiles come only from opt-in, paid peers bucketed by industry and region, never published below five participants, and never traceable to a workspace. The full data-handling summary lives on our security page.
Tunable figures on this page (engine roster, retention windows) live in code and ship with the product rather than in a PDF that drifts. Questions? Ask us anything about the method.

