Skip to main content
LLM Metrix
Back to tutorials

Brand Safety Monitoring Setup

Find the claims and reputationally loaded topics AI engines attach to your brand, judge how urgent each one is, and fix the source rather than arguing with the model.

Level

Intermediate

Format

Guide

Duration

9 min read

Sections

6 sections

Brand safety here means two signals on the same card: an AI engine stating something about you that is outdated, unverifiable or simply wrong, and an engine associating you with a reputationally loaded topic (a lawsuit, a scandal, a boycott) even when the statement is true. Neither is a keyword watchlist, and there is nothing to configure: the analysis runs on every scan, on every answer, whether or not you go looking for it.

What you configure is where the findings reach you. This tutorial covers both halves: reading the findings, and making sure the serious ones do not sit unread in a dashboard.

Step 1: Know what is measured on every answer

When an engine answers one of your prompts, the analyser records (alongside mention, rank and sentiment) two extra signals that exist purely for this:

  • Claims about your brand, plus an accuracy risk band. Specific factual assertions the answer makes: pricing, features, funding, ownership, comparisons. Statements that could be checked, quoted verbatim, and how likely they are to be outdated, unverifiable or inaccurate: none, low, medium or high.
  • Association topics, plus a controversy risk band. Free-text topics the answer ties to your brand that are reputationally loaded. A true-but-damaging association raises this band, not accuracy risk. Topics are not a category you can filter.

Both are per answer, so the same engine can be clean on one prompt and risky on the next, and two engines can disagree completely on the same question. One answer can raise both signals.

This is the product’s answer to a well-documented failure mode rather than a novel one. The survey literature on hallucination in large language models (see Huang et al. (2023)) distinguishes factuality failures, where a model asserts something untrue about the world, from faithfulness failures, where it departs from its source. What shows up in your brand’s answers is overwhelmingly the first kind, and it tends to be plausible, specific and confidently worded, which is exactly what makes it worth monitoring instead of assuming you would notice.

Step 2: Open the findings

Go to Alerts in the Monitor group of the sidebar. Below the alert list is a Brand Safety card, populated from your latest completed scan.

Findings are surfaced when the analyser saw a real risk, not for every answer. A low accuracy risk appears only if it also carried quotable claims; a low association risk appears only if it also carried at least one topic; medium and high always appear. A clean scan says so explicitly rather than showing an empty table.

The card sorts highest risk first, so the top row is where to start.

app.llmmetrix.com/dashboard/alerts
Illustrative, sample figures in the product's real layout

Brand Safety

AI engines sometimes state outdated claims or associate your brand with reputationally loaded topics. Review these.

These are heuristic flags for your team to review. Verified comparisons appear when ground-truth sources are attached.

AccuracyHigh riskGeminiwhen was Northwind founded and who funded it
  • Northwind was acquired by a larger analytics vendor in 2024.
  • Northwind raised a $40M Series B led by an unnamed investor.
AssociationHigh riskPerplexityis Northwind in legal trouble
  • class-action over billing
AccuracyMedium riskChatGPThow much does Northwind cost
  • Northwind starts at $19 per month for the entry tier.
AccuracyLow riskClaudewhat does Northwind do
  • Northwind is primarily a keyword rank tracker.
Live component, sample data. The Brand Safety card, sorted highest risk first. Accuracy and Association chips mark the two signals; a low-risk accuracy row appears only because it also carried quotable claims.

Step 3: Read a finding properly

Each finding is either Accuracy or Association. Both carry a risk badge, the engine that produced it, and the prompt that triggered it. Accuracy rows list the claims; Association rows list the topics. They are quoted, so you are reading what the model actually said rather than a paraphrase of it.

Read the prompt as carefully as the claim. A wrong statement produced by a hostile or leading question is a different problem from the same statement produced by “what does your brand do”, which is the question a genuine prospect asks. The second is far more urgent even at the same risk rating.

Note what is deliberately absent from this card: there is no safety score, no resolution workflow and no export. The findings are a read of one scan. Tracking a finding to closure is your process, not a field in the product, and the honest reason is that “resolved” is not something a scan can observe. What a scan can observe is whether the claim reappears in the next one.

What the card does offer is a way into remediation: each finding links to the published AI brand-safety remediation guide, or to a matching GEO recommendation from your latest scan when one exists: a related content gap or an engine-presence gap on the same engine. That is a link, not a workflow: the guide and the recommendations tell you what to publish and where; the product does not track the fix to closure.

app.llmmetrix.com/dashboard/alerts
Illustrative, sample figures in the product's real layout

Brand Safety

No inaccurate claims or reputationally loaded associations in the latest scan.

These are heuristic flags for your team to review. Verified comparisons appear when ground-truth sources are attached.

No accuracy or association findings in your latest scan.

Live component, sample data. The same card after a clean scan. It says so explicitly rather than showing an empty table: an empty table would read as a failure to load.

Step 4: Verify before you respond

The first move on an Accuracy finding is to check whether the claim is true. A meaningful share of them are, and reacting to a correct-but-unflattering statement as if it were misinformation wastes the week. An Association finding is the other case: the topic may be true, and the work is to review the answer, not to disprove a hallucination.

Work through it in this order:

  1. Is it true today? Pricing, positioning and feature claims are the most common findings and the most likely to be true-but-stale (right eighteen months ago, wrong now).
  2. Where could it have come from? Check the Citations page for the same scan. If the engine cited sources for that prompt, one of them is frequently the origin, and that is your fix.
  3. How many engines say it? One engine is noise; three engines converging on the same wrong statement means it is in material they all reach, which is a much bigger job and a much clearer one.
  4. Is it reachable by a buyer? A claim surfaced only by an obscure prompt is lower priority than the same claim on your main category question.

Record the prompt, the engine and the date alongside the quoted claim. Answers are generated fresh each time and are not guaranteed to reproduce, so the scan record is the evidence.

Step 5: Make sure findings reach you

Findings you never read are not monitoring. Two mechanisms carry them out of the dashboard:

  • Warning-level alerts. A medium- or high-risk Accuracy finding becomes a warning alert naming the engine and quoting the claim. A medium- or high-risk Association finding does the same and quotes a topic. So does an engine that describes your brand negatively overall, and an engine that failed to mention you on any tracked query.
  • Alert emails. With “Alert emails” on in Settings → Notifications, a scan that produces any warning-level alert emails you regardless of your digest cadence. That is the setting that matters here: someone on a monthly digest still hears about a hallucination the day it appears.

Outbound webhooks carry the same alert text if configured. One caveat that catches people: a competitor benchmark scan is silent on every outbound channel unless you have turned on “Notify on competitor scans” in Settings → Notifications. And it goes further than the notifications: the Alerts page, the Overview, the Brand Safety card, the shared report, the print view and the public API all read your brand’s own home-market scan, so a benchmark’s brand-safety findings are not sitting in the dashboard waiting to be noticed either. They are recorded in full on the scan; reading them means asking for that subject explicitly, with GET /api/v1/scans?project_id=…&subject_type=competitor&subject=Acme, where every answer comes back with its extracted claims and accuracy risk. Plan the benchmark around that, not around a card you will go looking for.

Setting Up Smart Alerts walks through the delivery settings in full.

app.llmmetrix.com/dashboard/settings?tab=notifications
Illustrative, sample figures in the product's real layout

Notification preferences

Email digest

Receive a visibility summary by email.

Alert emails

Email me when a scan surfaces something that needs attention.

In-app notifications

Show scan warnings, billing and delivery notices in the dashboard bell.

Pause alerts

While paused, per-scan alert emails, Slack and Teams stay quiet. The webhook and your scheduled digest still run. Handy while you're working through an incident and don't need the same finding interrupting you on every scan.

Receives a JSON POST on each completed scan. Payload reference.

Posts a summary when a scan surfaces something worth knowing.

Posts a message card when a scan surfaces something worth knowing.

Engines to alert about

Engine-specific findings (an engine that stopped mentioning you, a hallucination flag, negative tone) only reach email, webhook, Slack and Teams for the engines you leave checked. The Alerts page always shows every engine regardless; this is a delivery preference.

Competitor scans

Benchmark scans of competitor domains notify nobody by default: they are not news about your own brand. Agencies benchmarking a client's rivals can opt in here.

Alert types

Each type has a delivery mode: Immediate emails you when something needs attention, Digest reaches your digests and webhook without interrupting, Off silences it everywhere; the Alerts page always shows everything regardless. The daily, weekly and monthly digest follows the same modes.

AI Visibility Score change

When your score moves by ≥ 5 points since the previous scan

points
points (0 = off)

New citation detected

When an engine cites a source it did not cite in your previous scan, or stops citing one it did

Hallucination detected

When an engine states a claim about your brand that looks inaccurate, or describes you negatively

Competitor spike

When a competitor gains ≥ 10 points of share of voice since the previous scan

points

Sentiment shift

When sentiment falls by 0.15 on the 0–1 scale since the previous scan

on the 0–1 scale

Position change

When your average position in answers moves by ≥ 1 position, or your first-mention share drops by ≥ 10 points, since the previous scan

positions
points
Live component, sample data. Settings → Notifications with the digest off and alert emails on. That combination is the one this step argues for: the only mail you get needs a decision.

Step 6: Fix at the source, then re-measure

You cannot edit a model’s answer, and there is no channel for filing a correction. What you can change is the material an answer is built from.

  • Correct the third-party page. If a cited source carries the wrong figure, that is the highest- leverage fix available, and it is usually the fastest. A directory listing or review-site profile with stale data is often a form submission away.
  • Publish an unambiguous statement of the fact on your own site. One page, plainly worded, easy to retrieve. Vague marketing copy is exactly the input that produces a confident wrong answer.
  • Retire or update the stale source you control. An old pricing page or announcement that still ranks is a supply of the wrong answer.
  • Escalate as a communications issue when it is one. A reputational claim spreading across several engines is not an SEO ticket.

Then re-scan and check the same prompt. Treating this as a monitored loop rather than a one-off is the same posture the NIST AI Risk Management Framework asks of organisations deploying AI systems: identify, measure, and keep measuring, because the system’s behaviour changes underneath you. Here the system is somebody else’s, which makes the measuring the only lever you own.

Step 7: Where to go next

Ready to put this into practice?

Start optimizing your AI visibility with the techniques you've learned.