Same Prompt, Different Answer: How Many Runs Before You Trust the Number?
AI engines are non-deterministic, so a single scan is an anecdote with a decimal point. Here's roughly how many samples a mention rate needs before a change means anything — and why the honest answer is 'fewer than you'd like, more than you're running.'
Run the same prompt through the same engine twice and you will often get two different answers. Different wording, sometimes a different set of brands, occasionally a different recommendation — with nothing about your website, your competitors, or the web having changed in between.
This is not a defect. It is how these systems work, and every AI visibility number you look at is a sample from a distribution rather than a reading from an instrument. Most teams know this abstractly and then report the numbers as though it were not true.
The practical question is: how many samples before a difference means something?
Where the variance comes from
Three sources, with different implications.
Sampling during generation. Language models produce text by sampling from a probability distribution over next tokens. Temperature governs how much randomness that involves, and at anything above zero the same input can produce different outputs. This is deliberate — deterministic decoding produces noticeably worse and more repetitive text. The technique of sampling a model repeatedly and aggregating is well established in the literature; self-consistency demonstrated that sampling multiple reasoning paths and taking the majority substantially outperforms any single greedy output. The relevant point here is that variation across runs is the expected behaviour of these systems, not a symptom.
Retrieval variation. For grounded engines, each query triggers a live search. Results shift with index updates, freshness weighting, personalisation and load. Different documents retrieved means a different answer composed, and this source of variance is often larger than the sampling one.
Genuine change. Your content changed, a competitor published, the model was updated, the market moved. This is the only category you want to detect — and it arrives mixed in with the other two.
The whole measurement problem is separating the third from the first two.
What one scan actually tells you
Suppose you track 20 prompts across 5 engines. That is 100 observations, and you are mentioned in 40. Mention rate: 40%.
That figure has a confidence interval. For a proportion around 40% with 100 observations, the 95% interval is roughly ±10 percentage points — your true rate is somewhere in the vicinity of 30% to 50%.
Now next week’s scan reports 45%. Is that improvement?
No. It is entirely consistent with nothing having changed. The intervals overlap heavily, and you have observed one draw from each of two distributions that may well be identical. A team that reports “+5 points” here is reporting noise with a plus sign attached.
This is the single most common error in AI visibility reporting, and it is not a subtle one — it is reporting the difference between two samples as though it were a difference between two populations.
Roughly how many samples you need
The uncomfortable arithmetic, for detecting a change in a proportion near 50% (the worst case, where variance is highest):
| Observations | Approx. 95% margin | Smallest change you can call |
|---|---|---|
| 50 | ±14 pts | very large only |
| 100 | ±10 pts | ~20 pts |
| 400 | ±5 pts | ~10 pts |
| 1,000 | ±3 pts | ~6 pts |
| 2,500 | ±2 pts | ~4 pts |
Two things follow, and the second is the useful one.
Detecting small changes is expensive. Resolving a 5-point move reliably needs something in the low thousands of observations. At 5 engines that is hundreds of prompts, or dozens of prompts repeated many times — and every observation is a paid model call. This is a real constraint, not a tooling limitation.
But you rarely need to detect a 5-point move. The changes that matter operationally are large: an engine that stopped mentioning you entirely, a competitor that went from absent to dominant, a category where you fell out of the recommendation set. Those are 30-point moves and they are visible at modest sample sizes. Precision is expensive; usefulness is not.
Four ways to get more signal without more spend
Aggregate up, not down. Overall mention rate across all prompts and engines has your full sample behind it. Per-engine-per-prompt cells have a handful of observations each and are almost pure noise. Read the aggregate for trend; drop to the cell level only to investigate a specific question, and never report a single-cell change as news.
Use trend, not deltas. Six scans trending 38, 41, 39, 44, 43, 47 tell you something a single 38→47 comparison cannot, because the intermediate points constrain how much of the movement could be noise. Five points on a chart is worth more than one arrow.
Fix everything you can. The same prompts, the same engines, the same region, the same cadence. Variance you cannot eliminate is at least held constant, so a change in the output is more likely to reflect a change in the world. Rotating your prompt set and then comparing rates across the rotation compares two different measurements.
Treat position separately. Whether you were mentioned is binary and behaves well statistically. Where you appeared is ordinal and noisier still — position moves between runs constantly. Read position as a distribution over many observations rather than as a value; mentioned is not recommended covers what that distribution tells you.
What this means for cadence
Scan frequency is usually discussed as a freshness question. It is really a sample-accumulation question.
Daily scanning does not primarily give you fresher data — the underlying reality moves slower than that. It gives you more observations, which is what makes a genuine change distinguishable from resampling. Thirty daily scans a month produce a defensible trend; one monthly scan produces twelve isolated anecdotes a year.
That is the actual argument for cadence, and it is a better one than freshness. How often should I monitor AI visibility works through the cost trade-off, and multi-engine monitoring is where the accumulated series lives.
There is a corollary worth stating: a large one-off audit is worse value than a smaller recurring one. The audit gives you a wide snapshot with no baseline to compare against, which means its most confident-looking findings are the least verifiable.
The counter-argument
The reasonable objection: this is statistical over-engineering for a marketing function. Nobody runs power calculations on their content programme, and demanding confidence intervals before acting will paralyse a team that could have spent the time publishing.
Largely fair, and the practical concession is real. You do not need a statistics workflow. You need one habit: before reporting a change, ask whether it is bigger than the noise. In practice that reduces to a rule of thumb — a swing under about ten points on a hundred-ish observations is not news, and anything you would act on should be visible across several consecutive scans rather than one.
The reason to care at all is not rigour for its own sake. It is that acting on noise is expensive in a specific way: you attribute a random fluctuation to last month’s work, conclude the tactic worked, and scale it. Two months later you conclude AEO does not work when the same tactic “stops working.” Both conclusions came from the same non-signal, and the cost was a quarter of misdirected effort.
The summary
Every AI visibility figure is a sample. One scan is an anecdote, a few dozen prompts is a rough estimate, and detecting small changes reliably costs real money.
None of which means the numbers are useless — the large changes are the ones worth acting on, and they show up clearly at sample sizes anyone can afford. It means the small ones deserve silence rather than a slide. Why queries return different results covers the underlying non-determinism in more depth; the discipline it implies is simple enough to remember: report trends, not deltas, and let a change prove itself across more than one scan.
Written by
Team @ LLM MetrixWe research and write about AI brand visibility, GEO, AEO, and the evolving AI search landscape.
See how your brand appears in AI search
Track your visibility score across ChatGPT, Claude, Gemini, Perplexity, and more — free to start.
