Skip to main content
LLM Metrix
Safety · Compliance

Catch AI hallucinations
on every scan.

LLM Metrix flags potentially inaccurate claims an engine makes about you, plus the reputationally loaded topics it associates with your brand. These are heuristic flags for your team to review: analyzer likelihoods, not checks against a source of truth. Verified comparisons appear when ground-truth sources are attached.

Free forever plan. No card required

app.llmmetrix.com/dashboard/alerts
Illustrative, sample figures in the product's real layout
Brand Safety

Inaccurate claims and reputationally loaded topics, each with a risk band. Heuristic flags to review, not checks against a source of truth you upload.

High risk
2
Medium risk
1
Low risk
1
AccuracyHigh riskGeminihow much does Acme Corp cost
  • Acme Corp's enterprise plan starts at $1,200 per user per month.
AccuracyMedium riskChatGPTwho owns Acme Corp
  • Acme Corp was acquired by Google in 2023.
  • Acme Corp has around 40 employees.
AssociationHigh riskPerplexityis Acme Corp in legal trouble
  • class-action over billing
AccuracyLow riskClaudewhat does Acme Corp integrate with
  • Acme Corp only supports integrations with Slack.

Works with the AI engines your customers use

ChatGPTPerplexityGeminiClaudeGrokMeta AIDeepSeekGoogle AI OverviewsMicrosoft Copilot

Why it matters

Catch the wrong story on the scan that said it.

Hallucination detection

Automatically flag the factual claims an answer makes about you, pricing, features, funding, head-to-head comparisons, when they look inaccurate or risky, surfaced for your team to review.

Bands that set the order

Every finding carries a low, medium or high risk band, and the Alerts page lists them highest-band first, accuracy findings ahead of loaded associations, so triage starts at the riskiest statement instead of at a blank queue.

Per-engine risk view

See which engines the flagged claims come from, so you know where to focus.

Association detection

When an answer associates your brand with a reputationally loaded topic, a lawsuit, a scandal, a boycott, the topic and a risk band land on Alerts, separate from hallucination flags. Topics are free text, not a category you can filter.

Issue counts over time

High-, medium- and low-risk findings plotted scan by scan, so you can see whether flagged claims are clearing or accumulating. The wording of each claim belongs to the scan that found it, the trend counts them, the latest scan quotes them.

Sentiment monitoring

Every answer is scored for sentiment, plotted per scan on a 0–100 scale where 60 means every delivered answer reads neutral, alongside the positive, neutral and negative answer counts. You are told when engines describe your brand less positively than in the previous scan, reported as how many answers turned negative, not as a signed index.

Mechanics

How a misstatement gets caught, evidenced and tracked.

  1. 01

    Grade every delivered answer

    The analyzer pulls the factual statements each answer makes about your brand and grades the answer low, medium or high for accuracy risk. Low only counts when there are quotable statements to review.

  2. 02

    Flag loaded associations separately

    Reputationally loaded topics an answer attaches to your brand are recorded as free-text topics with their own risk band, kept apart from accuracy findings and alerted at medium and high.

  3. 03

    Record findings where history survives

    Every derived finding is recorded server-side, what fired, when, and on which project, with per-user read state, so two teammates reviewing the same incident never lose track of who has seen it.

  4. 04

    Re-check risky statements against your pages

    Attach pages on your own domain or a verified host under Settings → Ground truth and medium- and high-risk claims are re-checked after the scan; a match adds the quote and source link to the Alerts row, and a miss leaves the original band in place.

  5. 05

    Plot counts and mood per scan

    High-, medium- and low-risk counts plot beside the 0–100 sentiment line, giving direction over time without pretending history stores the wording, that stays readable on the scan that found it.

Who it's for

Who guards the brand

Comms lead catching AI misstatements before customers doThe wrong pricing story surfaces on your dashboard, not in a support ticket.

You own brand reputation and currently hear about AI errors only when a customer forwards a screenshot.

  • Findings arrive on Alerts with the engine and prompt that produced them, context enough to judge severity fast.
  • The alert quotes the riskiest graded statement, so triage starts from the words themselves.
  • Warning-level findings can reach email, Slack or Teams immediately; everything else waits for the digest cadence you set.
  1. 01Run a scan and read the flagged claims on the Alerts page.
  2. 02Attach your FAQ and pricing pages under Settings → Ground truth so medium- and high-risk statements get re-checked.
  3. 03Route warning-level findings to email or Slack immediately under Settings → Notifications.

Metrics this role tracks: High-risk accuracy findings · Association findings · Sentiment score

Review today's findings
Legal/compliance reviewer needing evidence not vibesEscalations carry a quote and a link, or say plainly they don't.

You sign off on external statements and need sourcing attached before anything leaves the building.

  • Attach pages on your own domain or a verified host; medium- and high-risk claims are re-checked against those pages after each scan.
  • When a match lands, the Alerts row carries the evidence quote and source URL alongside the original band.
  • A fetch or analyzer miss leaves the heuristic band untouched, flags never borrow certainty they don't have.
  • Associations stay topic-plus-band for human review; nothing auto-verifies them.
  1. 01Attach the pages that state your official facts under Settings → Ground truth, own domain or verified hosts only.
  2. 02Open the newest high-band finding on Alerts and read the quoted statement before escalating.
  3. 03Share outside counsel a token-gated report link rather than a seat.

Metrics this role tracks: High-risk accuracy findings · Medium-risk accuracy findings · Association findings

Attach ground-truth sources
Agency owner safeguarding client brandsEach client's risk picture lives in its own project.

You run retainers across several brands and need misstatement watch tuned per client without cross-talk between accounts.

  • Alert thresholds override per project, so a sensitive client runs stricter rules than the rest of the roster.
  • Alert history is recorded server-side per project, who saw what, and when, and survives team changes.
  • Hand clients a token-gated report link rather than seating them inside the workspace.
  1. 01Raise alert thresholds for the sensitive client from Settings → Notifications, scoped to that project alone.
  2. 02Run each client brand as its own project so findings never cross accounts.
  3. 03Send clients the token-gated report link instead of a workspace invite.

Metrics this role tracks: Accuracy findings (high / medium / low) · Association findings · Sentiment score

Tune per-project thresholds
Exec watching sentiment slide before revenue doesSentiment becomes a line on a chart, not a feeling in a meeting.

You track brand health monthly and want the AI-answer reading on the same dashboard as everything else you manage by.

  • A 0–100 sentiment line per completed scan, where 60 means every delivered answer read neutral.
  • The distribution card shows positive, neutral and negative answer counts right beside the line.
  • The drop alert reports how many answers turned negative, a count you can put in front of a board.
  1. 01Run a scan to plot your first point on the sentiment trend on the Alerts page.
  2. 02Switch on the sentiment-drop alert mode under Settings → Notifications.
  3. 03Put the 0–100 line beside its positive, neutral and negative counts in your monthly review.

Metrics this role tracks: Sentiment score · High-risk accuracy findings · Association findings

Open the sentiment trend
Support lead defusing the ticket spike a wrong answer causesThe misstatement surfaces here first, with the words attached.

You run support and hear about AI errors when tickets pile up, hours after the answer started spreading.

  • Accuracy findings arrive with the engine and the prompt that produced them, context enough to draft one correct reply and reuse it.
  • The alert quotes the riskiest graded statement, so macros answer what the engine actually said, not a retelling.
  • Findings sort highest-band first, so the worst offenders are written up before the marginal ones.
  • Warning-level findings can reach email, Slack or Teams at once; the rest follows the cadence you set.
  1. 01Attach your FAQ and pricing pages under Settings → Ground truth so risky claims are re-checked against them.
  2. 02Work the highest-band findings on the Alerts page into reply macros, quoting what the engine actually said.
  3. 03Switch warning-level findings to immediate delivery under Settings → Notifications.

Metrics this role tracks: Medium-risk accuracy findings · High-risk accuracy findings · Sentiment score

Open the findings queue

Foundations

How do you monitor AI hallucinations about your brand?

Every answer a scan delivers is graded for accuracy. An analyzer extracts the factual statements each engine makes about your brand, steered toward pricing, features, funding and head-to-head comparisons, and assigns the answer a low, medium or high risk band. The low band only counts when the answer also carries quotable claims, so the numbers reflect findings you could actually read and act on rather than answers that merely hedged.

These are heuristic flags for your team to review, not checks against a source of truth you upload. Attach pages under Settings → Ground truth and medium- and high-risk claims are re-checked after the scan; a miss leaves the original band in place.

Findings land on the Alerts page carrying the engine and the prompt that produced them, sorted highest-risk first. The wording of each flagged statement belongs to the scan that found it: later scans plot fresh counts and never rewrite what an earlier answer said, so a statement you are still investigating reads today exactly as the engine said it then.

  • Risk is graded per answer, and the alert quotes the riskiest graded statement.
  • Accuracy findings and loaded associations sort together, highest band first.
  • The trend is drawn from counts per scan; the wording stays with its own scan.

Two families

Is an inaccurate claim the same thing as a damaging association?

No, and keeping them apart is deliberate. An accuracy finding says an engine stated something false or risky about you, a wrong price, a feature you never shipped. An association finding says something different again: an answer connected your brand to a reputationally loaded topic such as legal trouble, a scandal or a boycott, whether or not anything inaccurate was asserted along the way. A true-but-damaging association never moves the accuracy band, which is exactly why it needs its own signal.

Association topics arrive as free-text strings plus a none-to-high risk band, not as fixed categories you can filter, the analyzer reports each topic the way the answer framed it, and a human decides what it means. Medium and high associations raise alerts on the same path as accuracy findings, so both families reach the same people through the same channels instead of one of them living only in the dashboard.

Mood, measured

How do you watch brand sentiment in AI answers?

Every delivered answer is scored positive, neutral or negative, and each scan freezes the numeric mean, which turns mood into a plottable 0–100 line where 60 means every delivered answer read neutral. A distribution card sits beside it showing the raw positive, neutral and negative answer counts per scan, so a slide is visible both as an average moving and as answers changing side.

When engines describe your brand less positively than in the previous scan, the alert says so in answer counts, how many turned negative, rather than as a signed index, because averaging categories has no natural zero and an index would imply precision the measurement does not have. Sentiment is also strictly brand-level: one value per answer about your brand, not a monitor on named individuals.

FAQ

Brand safety questions, answered.

What does brand safety monitoring catch?
Two things, on every scan. First, the factual claims an engine makes about your brand, pricing, features, funding, head-to-head comparisons, flagged with an accuracy-risk band when they look wrong or risky. Second, the reputationally loaded topics an answer associates with your brand, such as legal trouble or a boycott, each with its own risk band. Both kinds of finding land on the Alerts page with the engine and the answer they came from.
How do you know what counts as a hallucination?
Each answer is first flagged with a heuristic accuracy-risk band. If you attach pages on your own domain or a verified host in Settings → Ground truth, a follow-up check compares medium- and high-risk claims to those pages and, when it can, puts a quote and source link on the Alerts row. A fetch or analyzer miss leaves the heuristic flag in place, only those attached pages are fetched.
How often are claims checked?
Whenever a scan completes. Tracked prompts re-run automatically every day on paid plans and once a week on free, plus any scan you run yourself, and each completed scan re-grades the answers for accuracy risk and loaded associations.
Can I look back at findings from past scans?
You get counts per scan: high-, medium- and low-risk findings plotted over time, so you can see whether problems are clearing or accumulating. The wording of each individual claim belongs to the scan that found it, open that scan to read exactly what was said. Nothing is rewritten after the fact.
How is this different from automated alerts?
Brand safety monitoring produces the findings, flagged claims and risky associations with their risk bands, kept on your Alerts dashboard. Automated alerts are the delivery layer: score changes, competitor surges, new or lost citations, inaccurate claims, sentiment drops and position drift can each be sent to email, Slack, Microsoft Teams or an outbound webhook, immediately or via digest, each with its own switch. You are looking at the detector; alerts decide who hears about it.
Can you help us correct hallucinations?
Flagged claims sit alongside your GEO recommendations, which suggest the content and citation changes that can help correct how engines describe you. You act on them from your dashboard, we don't file takedowns or send outreach on your behalf.
Is this useful for non-public companies?
Absolutely, private companies are often hit hardest by hallucinations because there is less authoritative public data for engines to anchor on.
How do we get started?
Create a project and run a scan, flagged claims and associations appear on the Alerts page from the first completed scan. For stronger evidence, attach pages on your own domain or a verified host under Settings → Ground truth: after each scan, medium- and high-risk claims are re-checked against those pages, and when a match is found the quote and source link are added to the Alerts row. If a fetch misses, the original risk band simply stays in place.
Start on the free plan. No card, no time limit.

The AI era of search
is already here

See what AI engines say about your brand before your competitors do. Start free today. No card required.