Skip to main content
LLM Metrix
Back to tutorials

Citation Intelligence Deep Dive

Read the sources AI engines lean on for your queries, and understand why a citation with a real URL and a citation without one are two different measurements.

Level

Intermediate

Format

Guide

Duration

10 min read

Sections

6 sections

The Citations page answers a question no ranking view can: when an engine explains your category, whose material is it standing on? Those domains are the shortlist of places worth being present, and the absence of your own domain from them is usually a sharper finding than any score movement.

Before you act on it, though, you need to know which of two different things the page is showing you, because they look identical in the table and mean quite different things. That distinction is Step 2, and it is the most important part of this tutorial.

Step 1: Open Citations and read the four headline counts

Citations is in the Monitor group of the sidebar, subtitled “Pages cited by AI engines in their answers”. It reads your active project’s latest completed scan.

Four stat cards sit at the top: total citations recorded in that scan, unique domains among them, how many engines produced any citation at all, and the single most-cited domain with its count.

Take the third card seriously. If only two of your selected engines are citing anything, the rest of the page is a narrow sample, and Step 5 explains why that happens.

app.llmmetrix.com/dashboard/citations
Illustrative, sample figures in the product's real layout
5
Total Citations
latest scan
1
Your Domain Citations
latest scan
4
Domains cited
unique domains cited
3
URL-Backed Sources
retrieved source URLs
3
Engines Citing
engines that cited
Live component, sample data. Take the third card seriously: if only two engines cited anything, everything below it is a narrow sample.

Step 2: Check whether your citations carry real URLs

A citation on this page arrives by one of two routes, and they are not equally strong evidence.

Route one: a real cited source. The engine performed a web search as part of answering, and returned the URLs it used. That is a link the engine genuinely consulted. Perplexity’s search model does this on every call. Google AI Overviews carries Google’s own cited sources. ChatGPT, Gemini, Grok and Claude do it when search grounding is switched on for the deployment.

Route two: a domain extracted from prose. With grounding off, the engines are queried as plain chat models. No search happens, no links come back, and the analyser reads the answer text and pulls out the domains it names. The record it writes carries a domain and no URL and no title, because there was never a link to store.

Both appear in the table. You can tell them apart because a prose-extracted citation has nothing to click through to. Meta AI has no web-search capability at all, so its citations are always route two; a project scanned with grounding disabled will be entirely route two apart from Perplexity and, if configured, Google AI Overviews.

The interpretive difference is large. A real cited URL is evidence that a specific page was consulted while your category was being explained, the retrieval step that Lewis et al. (2020) formalised as retrieval-augmented generation, where a model’s answer is conditioned on documents fetched at query time. A prose-extracted domain is evidence only that the model named that domain, which may reflect its training data, a half-remembered association, or an outright confabulation.

So: use route-one citations to decide where to seek presence. Treat route-two domains as a weaker signal about which brands and publications the model associates with your category, and verify before you act.

app.llmmetrix.com/dashboard/citations
Illustrative, sample figures in the product's real layout
5 citations
2026-05-14
Query:best AI visibility tracking tool

northwind.co

2026-05-14
Query:northwind vs visify

softwareadvice.com

ChatGPT#1prose mention
2026-05-14
Query:best AI visibility tracking tool

wikipedia.org

Gemini#1prose mention
2026-05-14
Query:what is answer engine optimization

g2.com

2026-05-14
Query:is northwind any good

Links appear only for web-grounded scans; unlinked citations were extracted from answer text. Measured via API scans of your tracked prompts.

Live component, sample data. Both kinds of citation in one table. The rows with a title and an external-link icon were genuinely fetched; the bare domains were pulled out of prose and have nothing to click. The search box is live.

Step 3: Work the source-gap list

The Source gaps card is the part of the page to act on. It opens with a single line about your own domain (either a confirmation with a count, or an amber warning that no engine cited you at all) and then lists every third-party domain that was cited, ordered by how often, with a chip for each engine that cited it.

Read the own-domain line first. Not being cited while being mentioned is a specific and common condition: the engines know who you are and send readers somewhere else for the detail. That is a content-depth problem, not a brand-awareness one.

Then read the external list as a target list. These are the pages an engine reached for when answering your questions. Presence on them (a listing, a comparison entry, a review, a quoted contribution) is the mechanism by which you become part of the material the answer is built from.

Two cautions. First, the list is from one scan, so a single unusual answer can put a domain on it; look for domains that persist across scans and engines before spending on them. Second, a domain cited by every engine for every prompt is usually a general-purpose reference rather than an opportunity: being on Wikipedia is not the same kind of win as being in the category comparison that three engines keep quoting.

app.llmmetrix.com/dashboard/citations
Illustrative, sample figures in the product's real layout

Source gaps

External domains that appear in answers for your tracked queries. Start with the sources cited most often, then inspect the matching prompts below.

3 external domains
Your domain northwind.co was cited 1×.
×2

g2.com

Cited by 2 engines

91 DataForSEO RankPerplexityChatGPT
×1

softwareadvice.com

Cited by 1 engine

78 DataForSEO RankChatGPT
×1

wikipedia.org

Cited by 1 engine

99 DataForSEO RankGemini
Live component, sample data. Source gaps, leading with the own-domain line. Each third-party row carries a chip per engine that cited it, and a DA badge where an authority provider is configured.

Step 4: Use domain authority to order the list

When an authority provider is configured for the deployment, each external source carries a DA badge: a 0–100 domain authority figure from a nightly enrichment pass, cached per domain.

Use it for ordering, not as a verdict. A high-authority domain cited once is often a worse target than a mid-authority niche publication cited across four engines and six prompts, because the second one is demonstrably in the retrieval path for your questions. Citation frequency in your own scan is the primary sort; authority is the tie-break.

If no badges appear at all, no authority provider is configured: the enrichment simply has nothing to fill in, and nothing is broken.

Step 5: Compare which engines cite, and how much

The engine chips on each row let you see coverage per engine at a glance, and the search box above the citation table filters by domain, title, query or engine name.

Expect uneven results, and expect the unevenness to be structural rather than a fault:

  • Perplexity returns sources on every call by design.
  • Google AI Overviews is answered by a SERP provider rather than a chat model. If no SERP provider is configured for your deployment, this surface is skipped by scans entirely and will never contribute citations.
  • Meta AI has no search capability and will only ever produce prose-extracted domains.
  • ChatGPT, Gemini, Grok and Claude depend on whether grounding is enabled. Claude’s grounding needs a direct provider key, because its web-search tool is executed by the provider and cannot be passed through a gateway.

None of that is a ranking of the engines. It is a statement about how each was queried.

Step 6: Turn the list into outreach, not into on-site edits

The instinct after reading this page is to go and edit your own site. That is usually the wrong first move, because the finding is about other people’s pages.

  • Pursue presence on the sources that keep appearing. Listings, roundups, comparison pages and review sites in your category, in the order the page ranks them.
  • Make your own pages retrievable and quotable. A page that answers one question, states the answer plainly near the top, and is not blocked from AI crawlers can be used as a source; one that buries the answer under a narrative cannot. OpenAI documents its crawlers separately, and the one that feeds search results is not the same agent as the one that trains models. Blocking indiscriminately can remove you from the surface you are trying to appear on.
  • Correct what is wrong at the source. If a cited third-party page states something outdated about you, fixing that page changes what the engines have to work with. Editing your own site does not.
  • Re-scan and check the list moved. Citation change is slower than ranking change; give it more than a week.

Step 7: Where to go next

Ready to put this into practice?

Start optimizing your AI visibility with the techniques you've learned.