Catch AI hallucinations
on every scan.
LLM Metrix flags potentially inaccurate claims an engine makes about you, plus the reputationally loaded topics it associates with your brand. These are heuristic flags for your team to review: analyzer likelihoods, not checks against a source of truth. Verified comparisons appear when ground-truth sources are attached.
Free forever plan. No card required
Inaccurate claims and reputationally loaded topics, each with a risk band. Heuristic flags to review, not checks against a source of truth you upload.
- Acme Corp's enterprise plan starts at $1,200 per user per month.
- Acme Corp was acquired by Google in 2023.
- Acme Corp has around 40 employees.
- class-action over billing
- Acme Corp only supports integrations with Slack.
Works with the AI engines your customers use
Why it matters
Catch the wrong story on the scan that said it.

Hallucination detection
Automatically flag the factual claims an answer makes about you, pricing, features, funding, head-to-head comparisons, when they look inaccurate or risky, surfaced for your team to review.

Bands that set the order
Every finding carries a low, medium or high risk band, and the Alerts page lists them highest-band first, accuracy findings ahead of loaded associations, so triage starts at the riskiest statement instead of at a blank queue.

Per-engine risk view
See which engines the flagged claims come from, so you know where to focus.

Association detection
When an answer associates your brand with a reputationally loaded topic, a lawsuit, a scandal, a boycott, the topic and a risk band land on Alerts, separate from hallucination flags. Topics are free text, not a category you can filter.

Issue counts over time
High-, medium- and low-risk findings plotted scan by scan, so you can see whether flagged claims are clearing or accumulating. The wording of each claim belongs to the scan that found it, the trend counts them, the latest scan quotes them.

Sentiment monitoring
Every answer is scored for sentiment, plotted per scan on a 0–100 scale where 60 means every delivered answer reads neutral, alongside the positive, neutral and negative answer counts. You are told when engines describe your brand less positively than in the previous scan, reported as how many answers turned negative, not as a signed index.

Hallucination detection
Automatically flag the factual claims an answer makes about you, pricing, features, funding, head-to-head comparisons, when they look inaccurate or risky, surfaced for your team to review.

Bands that set the order
Every finding carries a low, medium or high risk band, and the Alerts page lists them highest-band first, accuracy findings ahead of loaded associations, so triage starts at the riskiest statement instead of at a blank queue.

Per-engine risk view
See which engines the flagged claims come from, so you know where to focus.

Association detection
When an answer associates your brand with a reputationally loaded topic, a lawsuit, a scandal, a boycott, the topic and a risk band land on Alerts, separate from hallucination flags. Topics are free text, not a category you can filter.

Issue counts over time
High-, medium- and low-risk findings plotted scan by scan, so you can see whether flagged claims are clearing or accumulating. The wording of each claim belongs to the scan that found it, the trend counts them, the latest scan quotes them.

Sentiment monitoring
Every answer is scored for sentiment, plotted per scan on a 0–100 scale where 60 means every delivered answer reads neutral, alongside the positive, neutral and negative answer counts. You are told when engines describe your brand less positively than in the previous scan, reported as how many answers turned negative, not as a signed index.
Mechanics
How a misstatement gets caught, evidenced and tracked.
- 01
Grade every delivered answer
The analyzer pulls the factual statements each answer makes about your brand and grades the answer low, medium or high for accuracy risk. Low only counts when there are quotable statements to review.
- 02
Flag loaded associations separately
Reputationally loaded topics an answer attaches to your brand are recorded as free-text topics with their own risk band, kept apart from accuracy findings and alerted at medium and high.
- 03
Record findings where history survives
Every derived finding is recorded server-side, what fired, when, and on which project, with per-user read state, so two teammates reviewing the same incident never lose track of who has seen it.
- 04
Re-check risky statements against your pages
Attach pages on your own domain or a verified host under Settings → Ground truth and medium- and high-risk claims are re-checked after the scan; a match adds the quote and source link to the Alerts row, and a miss leaves the original band in place.
- 05
Plot counts and mood per scan
High-, medium- and low-risk counts plot beside the 0–100 sentiment line, giving direction over time without pretending history stores the wording, that stays readable on the scan that found it.
Who it's for
Who guards the brand
Comms lead catching AI misstatements before customers doThe wrong pricing story surfaces on your dashboard, not in a support ticket.
You own brand reputation and currently hear about AI errors only when a customer forwards a screenshot.
- Findings arrive on Alerts with the engine and prompt that produced them, context enough to judge severity fast.
- The alert quotes the riskiest graded statement, so triage starts from the words themselves.
- Warning-level findings can reach email, Slack or Teams immediately; everything else waits for the digest cadence you set.
- 01Run a scan and read the flagged claims on the Alerts page.
- 02Attach your FAQ and pricing pages under Settings → Ground truth so medium- and high-risk statements get re-checked.
- 03Route warning-level findings to email or Slack immediately under Settings → Notifications.
Metrics this role tracks: High-risk accuracy findings · Association findings · Sentiment score
Review today's findingsLegal/compliance reviewer needing evidence not vibesEscalations carry a quote and a link, or say plainly they don't.
You sign off on external statements and need sourcing attached before anything leaves the building.
- Attach pages on your own domain or a verified host; medium- and high-risk claims are re-checked against those pages after each scan.
- When a match lands, the Alerts row carries the evidence quote and source URL alongside the original band.
- A fetch or analyzer miss leaves the heuristic band untouched, flags never borrow certainty they don't have.
- Associations stay topic-plus-band for human review; nothing auto-verifies them.
- 01Attach the pages that state your official facts under Settings → Ground truth, own domain or verified hosts only.
- 02Open the newest high-band finding on Alerts and read the quoted statement before escalating.
- 03Share outside counsel a token-gated report link rather than a seat.
Metrics this role tracks: High-risk accuracy findings · Medium-risk accuracy findings · Association findings
Attach ground-truth sourcesAgency owner safeguarding client brandsEach client's risk picture lives in its own project.
You run retainers across several brands and need misstatement watch tuned per client without cross-talk between accounts.
- Alert thresholds override per project, so a sensitive client runs stricter rules than the rest of the roster.
- Alert history is recorded server-side per project, who saw what, and when, and survives team changes.
- Hand clients a token-gated report link rather than seating them inside the workspace.
- 01Raise alert thresholds for the sensitive client from Settings → Notifications, scoped to that project alone.
- 02Run each client brand as its own project so findings never cross accounts.
- 03Send clients the token-gated report link instead of a workspace invite.
Metrics this role tracks: Accuracy findings (high / medium / low) · Association findings · Sentiment score
Tune per-project thresholdsExec watching sentiment slide before revenue doesSentiment becomes a line on a chart, not a feeling in a meeting.
You track brand health monthly and want the AI-answer reading on the same dashboard as everything else you manage by.
- A 0–100 sentiment line per completed scan, where 60 means every delivered answer read neutral.
- The distribution card shows positive, neutral and negative answer counts right beside the line.
- The drop alert reports how many answers turned negative, a count you can put in front of a board.
- 01Run a scan to plot your first point on the sentiment trend on the Alerts page.
- 02Switch on the sentiment-drop alert mode under Settings → Notifications.
- 03Put the 0–100 line beside its positive, neutral and negative counts in your monthly review.
Metrics this role tracks: Sentiment score · High-risk accuracy findings · Association findings
Open the sentiment trendSupport lead defusing the ticket spike a wrong answer causesThe misstatement surfaces here first, with the words attached.
You run support and hear about AI errors when tickets pile up, hours after the answer started spreading.
- Accuracy findings arrive with the engine and the prompt that produced them, context enough to draft one correct reply and reuse it.
- The alert quotes the riskiest graded statement, so macros answer what the engine actually said, not a retelling.
- Findings sort highest-band first, so the worst offenders are written up before the marginal ones.
- Warning-level findings can reach email, Slack or Teams at once; the rest follows the cadence you set.
- 01Attach your FAQ and pricing pages under Settings → Ground truth so risky claims are re-checked against them.
- 02Work the highest-band findings on the Alerts page into reply macros, quoting what the engine actually said.
- 03Switch warning-level findings to immediate delivery under Settings → Notifications.
Metrics this role tracks: Medium-risk accuracy findings · High-risk accuracy findings · Sentiment score
Open the findings queueFoundations
How do you monitor AI hallucinations about your brand?
Every answer a scan delivers is graded for accuracy. An analyzer extracts the factual statements each engine makes about your brand, steered toward pricing, features, funding and head-to-head comparisons, and assigns the answer a low, medium or high risk band. The low band only counts when the answer also carries quotable claims, so the numbers reflect findings you could actually read and act on rather than answers that merely hedged.
These are heuristic flags for your team to review, not checks against a source of truth you upload. Attach pages under Settings → Ground truth and medium- and high-risk claims are re-checked after the scan; a miss leaves the original band in place.
Findings land on the Alerts page carrying the engine and the prompt that produced them, sorted highest-risk first. The wording of each flagged statement belongs to the scan that found it: later scans plot fresh counts and never rewrite what an earlier answer said, so a statement you are still investigating reads today exactly as the engine said it then.
- Risk is graded per answer, and the alert quotes the riskiest graded statement.
- Accuracy findings and loaded associations sort together, highest band first.
- The trend is drawn from counts per scan; the wording stays with its own scan.
Two families
Is an inaccurate claim the same thing as a damaging association?
No, and keeping them apart is deliberate. An accuracy finding says an engine stated something false or risky about you, a wrong price, a feature you never shipped. An association finding says something different again: an answer connected your brand to a reputationally loaded topic such as legal trouble, a scandal or a boycott, whether or not anything inaccurate was asserted along the way. A true-but-damaging association never moves the accuracy band, which is exactly why it needs its own signal.
Association topics arrive as free-text strings plus a none-to-high risk band, not as fixed categories you can filter, the analyzer reports each topic the way the answer framed it, and a human decides what it means. Medium and high associations raise alerts on the same path as accuracy findings, so both families reach the same people through the same channels instead of one of them living only in the dashboard.
Mood, measured
How do you watch brand sentiment in AI answers?
Every delivered answer is scored positive, neutral or negative, and each scan freezes the numeric mean, which turns mood into a plottable 0–100 line where 60 means every delivered answer read neutral. A distribution card sits beside it showing the raw positive, neutral and negative answer counts per scan, so a slide is visible both as an average moving and as answers changing side.
When engines describe your brand less positively than in the previous scan, the alert says so in answer counts, how many turned negative, rather than as a signed index, because averaging categories has no natural zero and an index would imply precision the measurement does not have. Sentiment is also strictly brand-level: one value per answer about your brand, not a monitor on named individuals.
FAQ
Brand safety questions, answered.
What does brand safety monitoring catch?
How do you know what counts as a hallucination?
How often are claims checked?
Can I look back at findings from past scans?
How is this different from automated alerts?
Can you help us correct hallucinations?
Is this useful for non-public companies?
How do we get started?
Keep exploring
Related features
The AI era of search
is already here
See what AI engines say about your brand before your competitors do. Start free today. No card required.

