The visibility score is the number the whole dashboard hangs off, so it is worth knowing exactly what it is made of. It is not a ranking, not a percentage of anything, and not comparable to anyone else’s score. It is a composite index computed from your own scan, and every input to it is visible to you on other pages.
This tutorial follows one number all the way down: from a single AI answer, to a per-engine score, to the 0–100 figure at the top of the Overview page — and then back up, to work out which part of it to attack.
Step 1: Locate the score and the scan that produced it
Open Overview in the Monitor group of the sidebar. The headline score belongs to your latest completed scan for the active project; the trend chart below it plots the same figure across your scan history.
Two things follow from that, and both matter:
- The score is frozen at write time. It is computed once, when the scan is committed, and stored on the scan row. Re-opening the page a week later does not recompute it, so a number you screenshotted will still be the number in the dashboard.
- Before your first completed scan the dashboard shows sample data, clearly labelled as such. A score on a project that has never scanned is illustration, not measurement.
AI Visibility Score
Score over time
Step 2: Break one answer into the three things that get measured
Each engine is asked each of your prompts, and every answer that comes back is analysed for exactly three things:
- Was the brand mentioned at all? A yes or a no.
- How prominently? A 1-based rank within the answer — first named, second, and so on — or no rank at all when the brand appears without any orderable position.
- How was it described? Positive, neutral or negative.
Nothing else feeds the score. Citations, competitor mentions and claims are recorded on the same answer and drive other pages, but they are not score inputs.
Answers whose engine call failed are excluded from all of it. A timed-out request is missing data, not an absence of your brand, and counting it as “not mentioned” would let a flaky provider manufacture a visibility drop you would then go and try to fix.
Step 3: Follow the maths for one engine
A single engine’s 0–100 score is a weighted blend of those three components:
- Mention rate — half the score. Answers that mention you, divided by answers that came back.
- Prominence — just over a third. Each mention’s rank becomes a 0–1 factor: first place is worth the full amount, second and third progressively less, and anything past the middle of a list is worth a fraction of first. A mention with no rank contributes nothing here.
- Sentiment — the remainder. Positive scores full marks, neutral scores a majority of them, negative scores zero.
The subtlety that trips people up is the denominator. Mention rate is averaged over every answer the engine returned. Prominence and sentiment are averaged only over the answers that mentioned you — because an answer that never names you has no position and no sentiment to measure.
That produces a genuinely counter-intuitive result, and it is the single most useful thing to understand about this number. A brand named in one answer out of ten, first and positively, can score the same as a brand named in every answer, mid-list, with neutral framing. The first brand looks excellent on two components measured over a tiny sample; the second is everywhere and never recommended. Same score, opposite problems, opposite fixes. Always read the mention rate next to the score, never the score alone.
One more guard: if an engine mentioned you in nothing at all, the sentiment term is forced to zero rather than defaulting to neutral. Without that, a brand no engine has ever heard of would collect points for the sentiment of mentions that do not exist.
Step 4: See how per-engine scores become one number
The headline score is the mean of the per-engine scores across engines that returned at least one usable answer, weighted by how many answers each engine actually delivered.
The weighting is deliberate and worth knowing: an engine that answered two prompts counts a fifth as much as one that answered ten. It used to be an unweighted mean, and that was a defect rather than a design — a single flaky engine that limped in with one answer out of twenty carried the same authority as one that answered every prompt, which was enough to move the headline number by tens of points. That number is frozen onto the scan row, so it drives the trend line, the competitor ranking and the high-score alert; a partial outage could shift all three and nothing said so.
An engine whose every call failed is dropped from the average entirely rather than averaged in as a zero — its row would be all zeros, which is indistinguishable from a brand nobody mentions.
If no engine returned a single usable answer, no scan is written at all. The job fails and retries, because that is an infrastructure outage rather than a measurement, and committing a “0/100” scan for it would both mislead you and charge your credit ledger for nothing.
Step 5: Interpret the number without inventing a benchmark
There is no published industry average for this score, and you should be sceptical of anyone who quotes one. The figure depends entirely on which prompts you chose and which engines you selected, so two brands’ scores are comparable only if both were measured the same way.
What the number is comparable to is arithmetic, so reason from that instead:
- Mentioned in every answer, always first, always positively → 100.
- Mentioned in every answer, mid-list, neutrally → the low 70s.
- Mentioned in half of the answers, around third place, neutrally → the mid-50s.
- Never mentioned → 0, exactly.
Compare against your own history, and against competitors measured on the same prompt set — which is what the benchmark table on the Competitors page is for. A month-over-month move on your own prompts is real information; a comparison against a number on a vendor’s marketing page is not.
Expect some run-to-run wobble even when nothing changed on your side. These are sampling-based models, and the research on self-consistency methods (Wang et al., 2022) exists precisely because one sample from a language model is not a stable measurement. Read the trend line, not the last point.
Step 6: Diagnose which component is costing you
Because the components are separable, so is the fix. Work out which one is low before you write a word of content:
- Low mention rate, good prominence. The engines know you, but surface you for a narrow slice of questions. That is a coverage problem — use Content Gaps to find the prompts you are absent from, and widen the topics you publish on.
- Good mention rate, low prominence. You are in the list and never at the top of it. That is a positioning problem, and Answer Engine Rankings is the page that shows where in each answer you actually land.
- Negative sentiment. Rarer, and more urgent than either. Check Alerts for brand-safety findings; an engine describing you badly is usually repeating something specific and correctable.
- One engine far below the others. Open the per-engine breakdown before concluding anything. Engines differ in how they answer conversational, intent-bearing questions — Google’s own guidance on AI features describes surfaces that answer rather than rank — so an outlier is often a real per-engine difference rather than a fault on your side.
Engine Breakdown
View allStep 7: Where to go next
- Visibility score — the feature page behind this number.
- What is a good visibility score? — how to set a target that means something in your market.
- Understanding your visibility score — the conceptual companion to this walkthrough.
- Setting Up Tracked Prompts — the prompt set is the measurement instrument. Change it and the score is measuring something else.
