Why We Built LLM Metrix
AI is becoming a primary channel for brand discovery, and most marketing teams have no idea how they show up there. Here's the problem we set out to solve, the design decisions that fell out of it, and what we deliberately refuse to claim.
It Started with a Spreadsheet
About eighteen months ago, we were helping a friend who runs a B2B SaaS company think through why new customer acquisition felt like it had stalled. Organic search traffic was flat, paid channels were performing normally, and the product hadn’t changed.
While talking it through, something odd came up: a couple of recent customers had mentioned hearing about the company from “asking ChatGPT.” So we started asking new customers how they had first come across the product. The answers were mostly the traditional channels, as you would expect — but AI engines came up more than once, and more than they had a year earlier.
Then came the obvious follow-up question: what were those people told when they asked?
Nobody knew. There was no way to know. You could go ask ChatGPT right now, but that gets you a single data point from a non-deterministic system. What was the typical response? How was the brand described relative to competitors? Was the description even accurate? Was it improving or degrading?
So we built a spreadsheet. Every Monday morning, we manually queried three engines with fifteen category-relevant questions and recorded the responses. It took about two hours a week. After a month there were patterns worth looking at — an engine that never mentioned the brand at all, a competitor that showed up in nearly every comparison answer, a product description that was two positioning cycles out of date.
Two hours a week is a bad way to run a measurement program. It is manual, it does not scale past one brand, and it silently stops happening the first busy week. But it was enough to establish that the underlying thing — a systematic, repeated, comparable read on what AI engines say about a brand — was both possible and completely unavailable as a product.
The Problem We’re Solving
The simplest framing: AI engines have become a meaningful channel for brand discovery, and very few teams are measuring them.
This is not a speculative future problem. Pew Research Center’s June 2026 survey of 5,119 U.S. adults found that about half of American adults now use AI chatbots, roughly one in four of them daily, with searching for information the single most common use — reported by 42%. Six in ten said they read AI summaries in search results.
When someone in that group asks an engine for a recommendation in your category, the engine synthesizes an answer from whatever it knows about the brands in that space. If your brand is absent from that answer, described inaccurately, or mentioned fifth instead of first, you lose the consideration set silently. There is no traffic dip to investigate, because there was never a click to lose.
Traditional analytics were not built to see this. It is worth being precise rather than sweeping here, because the incumbents are moving: Google has introduced generative AI performance reporting in Search Console, which is a fair signal that the surface is real enough for Google to instrument. But reporting on your own performance inside Google’s AI features is a different question from what half a dozen independent engines say about you when someone asks a question in your category — and nothing in the conventional stack answers that one.
What We Built
LLM Metrix does three things.
It measures. We run your tracked prompts against seven surfaces — ChatGPT, Perplexity, Gemini, Claude, Grok, Meta AI and Google AI Overviews — on a schedule set by your plan, and analyse each answer: does your brand appear, in what position, with what sentiment, cited from which sources, alongside which of your tracked competitors. Three of those — appearance, position and sentiment — roll up into a single visibility score so there is one number to watch, with the per-engine and per-prompt breakdown underneath it. Citations and competitors are measured and reported on their own surfaces and carry no weight in that number, which is a deliberate separation rather than an omission: folding them in would let a strong week on one hide a bad week on the other, and a headline score that moves when a rival publishes something is a number about them.
It explains. A score without context is useless and slightly dangerous, because a number people cannot interrogate is a number they eventually invent explanations for. So every score decomposes: which prompts you appear in and which you do not, how each engine describes you, which sources each engine named instead of you, and where competitors are winning.
It alerts. When an engine stops mentioning you, or starts describing you negatively, or asserts something inaccurate, you hear about it — when the scan that found it completes, which makes your detection latency your scan cadence rather than zero. That cadence is weekly on every paid plan, and manual on free. Deliberately, most events do not alert either: a system that emails you about everything trains you to ignore it. Alerts covers where we drew that line and why.
The design decisions worth explaining
Three choices shaped the product more than anything on the feature list.
Measure the panel, not one engine. Engines disagree, and the disagreement is the most useful signal in the dataset. An engine that mentions you while five others do not is telling you something specific — usually that one source it leans on describes you well, and the rest have nothing to work with. A single-engine tool cannot surface that, and a blended score with no breakdown hides it.
Join it to what you already rank for. The most actionable output we ship is not an AI metric at all. The Search↔AI gap puts Google Search Console queries next to scanned prompts and reports where you rank on page one of Google but no engine mentions you for the matching question. Those are pages where crawlability, authority and topical fit are already proven, so the remaining fix is structural and cheap. Neither half of the data produces that view alone, which is exactly why nobody was producing it.
Refuse to overclaim. This one turned out to be the hardest, and it is the design principle we are proudest of.
What we deliberately do not tell you
An AI visibility tool sits on data with real gaps in it, and the tempting move — the one that makes for a much more impressive dashboard — is to fill those gaps with confident-sounding inference. We decided early not to.
Concretely: in the Search↔AI gap report, a Google query that matches none of your tracked prompts is not reported as “AI ignores you.” Nobody asked. And a prompt where an engine mentioned you but no matching search query exists is not reported as “you don’t rank” — Google withholds queries issued by only a handful of users, a behaviour it documents in its performance data deep dive, so an absent query is frequently a privacy filter rather than a ranking fact. Both categories are counted and shown as coverage, never asserted as findings.
The same restraint runs through the citation data. When a scan runs without web-search grounding there is no retrieval step and therefore no link — the analyzer extracts a domain from prose rather than following a URL. So URL-level citation analysis is switched off for those scans rather than run on empty data. Without that guard, zero citation URLs would make every page you rank for look “not cited”: a confident, wrong, and extremely plausible-looking finding that could go unchallenged for months.
That principle extends to how we talk about the product, including in this post. We have no published case study, no customer benchmark study, and no aggregate dataset we are willing to make claims from. So you will not find a “brands using LLM Metrix see N% more citations” line anywhere in our marketing, because we would have to invent it. A stated gap is more useful than a confident number with nothing behind it — and it is the same standard we hold the product’s own reports to. It would be incoherent to ship a tool that refuses to overclaim and then market it by overclaiming.
Where We’re Headed
We are still early and we know it. The AI search landscape is moving fast: engines launch, retrieval behaviour changes, and the line between search and assistant keeps blurring. We are building to absorb those changes rather than reactively catch up to them, which is mostly a statement about architecture — engine coverage is configuration rather than a rewrite, and the derivations are computed from stored answers rather than baked in at collection time, so a new way of reading the data does not require re-collecting it.
The direction we care most about is deepening the diagnostic half. Knowing you are invisible for a prompt is a start. Knowing which source the engine named instead, whether you rank for the corresponding search query, and what specifically about your page made it unquotable — that is the part that turns a dashboard into a work queue. A metric that goes down and gives you nothing to do about it is a worse product than no metric.
The longer-term bet is that AI search becomes as measurable and improvable as traditional search did, and that whoever builds the measurement layer helps define what “improving” even means. We would rather that definition rest on honest data with its gaps labelled than on impressive numbers nobody can check.
If you’re starting from zero
If you are reading this because you are trying to work out how your brand shows up in AI search, you are thinking about the right problem at the right time — and the first step costs nothing.
Write down the fifteen questions your buyers actually ask, including the comparison and alternatives-to ones. Ask them across a few engines in a clean, logged-out session. Record whether you appear, where, how accurately, and who appears instead. Then do it again in a month and compare.
That was our spreadsheet. It is still a reasonable place to start, and if it tells you everything you need, you do not need us. Most people find that the second month is the one where it stops happening — which is roughly the point at which automating it starts to matter.
Written by
Team @ LLM MetrixWe research and write about AI brand visibility, GEO, AEO, and the evolving AI search landscape.
More in Company
See how your brand appears in AI search
Track your visibility score across ChatGPT, Claude, Gemini, Perplexity, and more — free to start.
