Introducing Multi-Engine Visibility Scoring
We rebuilt our scoring engine from scratch to give you a single, unified visibility score that reflects your brand's presence across ChatGPT, Claude, Gemini, Perplexity, and more.
The Problem with Per-Engine Scores
When we launched LLM Metrix, we gave users a separate score for each AI engine: your ChatGPT score, your Claude score, your Perplexity score. It felt like the right call — each engine is a different product with a different audience, and showing them separately seemed transparent.
In practice it created more confusion than clarity. The feedback was consistent: “I have six numbers. I don’t know if I’m doing well.” A score of 71 on ChatGPT and 43 on Gemini doesn’t tell you whether your brand is visible. It tells you that you have work to do on Gemini, which you probably already suspected.
The deeper problem was that per-engine scores made progress impossible to read at a glance. If ChatGPT improved by 8 points and Gemini dropped by 4, what happened to your brand? You’d have to do the arithmetic yourself, and decide how to weight the engines while doing it. Nobody was doing that, which meant the numbers were being compared by eye — the least reliable method available.
Why a composite score is necessary at all
It’s worth saying why this needs inventing rather than borrowing. Classic search has a natural scalar: position. A generated answer doesn’t. There’s no ranked list to read a number off, so any single figure has to be constructed from the properties of the answer text itself.
The research literature hit the same wall. The 2023 paper GEO: Generative Engine Optimization had to define its own impression metrics for generative engines — weighting a source by how prominently it features in the generated text — rather than reuse click-and-rank measures. Our score applies the same idea to a brand rather than a URL.
What We Built
Every LLM Metrix account now shows a single Unified Visibility Score — one number from 0 to 100 representing your brand’s presence across the AI engines being monitored. Visibility score covers how to read it; this post is about how it’s built and why it’s built that way.
How the Score Is Calculated
Within each engine, three components:
Mention rate (50%) — how often the engine names your brand across the prompts you track. A brand that isn’t mentioned doesn’t score on this component.
Position (35%) — when your brand is mentioned, how prominently it appears. An earlier, more emphatic mention weighs more than a passing reference at the end of a list.
Sentiment (15%) — how your brand is described. Positive framing lifts the score, neutral holds steady, negative pulls it down. An AI classifier tuned for brand sentiment in engine responses does the labelling.
Each engine produces a 0–100 score from those three components, and your Unified Score is the average across engines.
Three design decisions worth explaining
A composite score is mostly a set of judgement calls, and the calls are more useful to you than the formula.
Engines are weighted equally
Your standing on Claude counts the same as your standing on ChatGPT. We considered weighting by market share and rejected it for two reasons. First, reliable, current usage share by engine is not something we can source honestly, and a weighting built on a number we made up would propagate into every customer’s headline metric. Second, market share is the wrong variable even if you had it: what matters is where your buyers are, which varies enormously by category and is not something a platform-wide constant can encode.
So the score answers “how present are you across the AI landscape?” If one engine matters disproportionately for your audience, the Engine Breakdown panel is where you should be looking, and it sits directly below the score for that reason. See also which AI engine matters most.
An engine that returned nothing is excluded, not scored zero
If every call to an engine fails — a rate limit, a provider outage, a timeout — that engine is dropped from the average rather than contributing a zero.
The reason is that a 0/100, 0%-mention-rate result is indistinguishable from genuine invisibility once it’s written down. Counting an outage as absence would mean a provider having a bad afternoon shows up in your dashboard as your brand losing visibility, and you would go looking for a content problem that doesn’t exist.
The same logic runs one level up. If no engine returns usable data, we don’t record a scan with a score of zero — we fail the job and retry it. A committed zero is a permanent, wrong data point in your history; a failed job is an inconvenience.
Sentiment is gated on there being at least one mention
This one came out of a bug worth describing, because it illustrates how a reasonable-looking default corrupts a metric.
Sentiment defaulted to a neutral value when there was nothing to measure. For a brand mentioned by no engine at all, that neutral default still flowed into the 15% sentiment term — so a brand that nothing had ever said anything about received a non-zero score. Not much: a floor of around nine points out of a hundred. But it was nine points of reputation credit for a brand with no presence to have a reputation about, and worse, it made the difference between “invisible” and “barely visible” impossible to see.
Sentiment is now gated on at least one mention existing. With no mentions, the term contributes nothing and the score is zero — which is the honest answer.
What Changed in the Dashboard
The per-engine bar chart is still there. It’s now the Engine Breakdown panel, sitting below the main score card, and you can still drill into per-engine trends from it.
The main score card leads with your Unified Score and its trend line, giving the at-a-glance read most people wanted while keeping the detailed per-engine data one click away.
How to use it — and how not to
Do track the trend rather than the absolute figure. Generated answers are sampled, so any single scan carries genuine variance. A score that has climbed steadily over a month is a signal; a four-point move week to week usually isn’t.
Do open the Engine Breakdown before drawing conclusions. A flat Unified Score can hide two engines moving hard in opposite directions.
Don’t compare your score against another company’s, or against a number you saw quoted somewhere. The score is computed over your tracked prompts. Two brands with different query sets have different denominators, and the figures aren’t comparable. What is comparable is you against yourself over time, and you against competitors measured on the same query set — which is what competitor benchmarking is for.
Don’t read the score as a market-share estimate. It measures presence in a defined set of answers, not commercial outcomes. See what is a good visibility score for how to set a realistic target.
The limits of a single number
Any composite hides its components — that’s what makes it useful and what makes it dangerous. Ours weights three properties in a fixed ratio that we chose, and reasonable people would choose differently. It cannot tell you why a number moved, and it cannot distinguish a genuine change in how engines describe you from a change in how they sampled this week.
That’s why the breakdown is one click away rather than behind a menu, and why every component is documented rather than described as proprietary. A score you can’t decompose is a score you can’t argue with — and if you’re going to put a number in front of your executive team, you should be able to explain exactly where it came from.
The Unified Score is live for all accounts. If you have questions about the methodology, reach out — we read every message.
Written by
Team @ LLM MetrixWe research and write about AI brand visibility, GEO, AEO, and the evolving AI search landscape.
See how your brand appears in AI search
Track your visibility score across ChatGPT, Claude, Gemini, Perplexity, and more — free to start.
