Which AI reads your site,
and which sends people.
Two halves of the same question. AI crawlers hit your pages to train models, build answer indexes or answer someone's live question; AI answer surfaces then send visitors back. This surface measures both halves of your AI traffic, and the ratio between them.
Free forever plan. No card required
Works with the AI engines your customers use
Why it matters
The crawlers and referrals classic analytics misses.

A crawler taxonomy, not a user-agent dump
AI crawler detection starts with classification: crawlers from OpenAI, Anthropic, Perplexity, Google, Apple, Amazon, ByteDance, Meta and Common Crawl, each tagged by operator and by why it visits, model training, answer-index building, or a real-time fetch for a live question. A hit reads as an operator and an intent rather than an opaque string.

Verification, because a user agent is free text
Anyone can claim to be GPTBot. Requests are checked against each operator's published IP-range file, or reverse-resolved and forward-confirmed the way Googlebot verification works. Operators publishing no mechanism stay honestly marked unverified. The verified figure is reported as a share of the total rather than replacing it, unconfirmed is not the same as fake, and only you can decide which to act on.

Referral attribution that survives stripped referrers
ChatGPT, Perplexity, Gemini, Claude, Microsoft Copilot and Meta AI are matched by referrer host or by a UTM tag you seeded, and the explicit tag wins, because AI apps routinely strip the referrer entirely and those ChatGPT referrals would otherwise look like direct traffic. A link builder generates the tagged URLs.

Daily rollups and a crawl-to-refer ratio
Crawls, unique paths and operator-attributed referrals per bot per day. A ratio over zero referrals reads as "no referrals yet", never as a plausible-looking zero, and an operator's referrals are counted once per day, not once per bot.

Crawled but never referred
The pages AI crawlers read that no AI referral ever landed on, the content the engines consume without sending anyone back. That list is the optimization backlog.

Two ingestion paths, each fenced
A first-party edge beacon for referrals, and server access-log upload for crawler hits. The beacon key is write-only and bound to one project; log upload additionally requires proven domain ownership.

A crawler taxonomy, not a user-agent dump
AI crawler detection starts with classification: crawlers from OpenAI, Anthropic, Perplexity, Google, Apple, Amazon, ByteDance, Meta and Common Crawl, each tagged by operator and by why it visits, model training, answer-index building, or a real-time fetch for a live question. A hit reads as an operator and an intent rather than an opaque string.

Verification, because a user agent is free text
Anyone can claim to be GPTBot. Requests are checked against each operator's published IP-range file, or reverse-resolved and forward-confirmed the way Googlebot verification works. Operators publishing no mechanism stay honestly marked unverified. The verified figure is reported as a share of the total rather than replacing it, unconfirmed is not the same as fake, and only you can decide which to act on.

Referral attribution that survives stripped referrers
ChatGPT, Perplexity, Gemini, Claude, Microsoft Copilot and Meta AI are matched by referrer host or by a UTM tag you seeded, and the explicit tag wins, because AI apps routinely strip the referrer entirely and those ChatGPT referrals would otherwise look like direct traffic. A link builder generates the tagged URLs.

Daily rollups and a crawl-to-refer ratio
Crawls, unique paths and operator-attributed referrals per bot per day. A ratio over zero referrals reads as "no referrals yet", never as a plausible-looking zero, and an operator's referrals are counted once per day, not once per bot.

Crawled but never referred
The pages AI crawlers read that no AI referral ever landed on, the content the engines consume without sending anyone back. That list is the optimization backlog.

Two ingestion paths, each fenced
A first-party edge beacon for referrals, and server access-log upload for crawler hits. The beacon key is write-only and bound to one project; log upload additionally requires proven domain ownership.
Mechanics
How it works, end to end.
- 01
Drop the beacon, seed the tags
Paste the beacon snippet into your site directly or through a tag manager, and mint tagged URLs with the built-in link builder. Each visit is classified by its utm_source first and its referrer host second.
- 02
Feed the crawler half
Upload server access logs, after proving domain ownership at the meta-tag tier, or connect Cloudflare. Every line's user agent is parsed against the taxonomy; human and non-AI traffic is set aside.
- 03
Classify operator and intent
Each hit resolves to the most specific matching signature, so a Claude UA carrying several tokens still lands on the right bot, and the row records who operates it and why it visits, training, index-building or a live fetch.
- 04
Verify the client IP
Where the operator publishes an IP-range file the request IP is tested against it; where reverse DNS is documented, the hostname is forward-confirmed. Any failure records unverified, nothing is dropped.
- 05
Roll up daily and read both halves together
Crawls, unique paths and verified shares roll up per bot per day, feeding the crawl-to-refer ratio and the crawled-but-never-referred list beside what the engines actually say about you in scans.
Who it's for
Who watches the bots
Technical SEO measuring crawler fetch patternsWhich bots hit which paths, how often, and whether they were who they said.
You manage crawl budget and robots policy and want per-operator evidence instead of a user-agent filter in a log viewer.
- The crawler table on the Traffic page groups hits by operator and purpose, training, index-building, live fetch, not opaque strings.
- Unique paths and daily rollups show where each operator spends its fetches across your site.
- A verified share sits beside every crawl total, telling you how much of each bot's traffic was confirmable.
- 01Verify your domain at the meta-tag tier from the Domains page, log upload requires it.
- 02Upload last month's access logs from the Traffic page.
- 03Group the crawler table by operator and intent to see who fetches what.
Metrics this role tracks: Verified crawler hits · Crawl-to-refer ratio · Referral sessions
Open AI TrafficContent ops justifying AI-era server costsSomeone has to pay for the crawl wave; show them exactly who causes it.
You own infrastructure budgets and need per-operator crawl volumes to explain the load AI crawlers add.
- Daily rollups quantify fetches per operator per bot, so cost conversations name culprits rather than vague trends.
- Heavy crawling reads as intent-tagged behaviour, a training crawler's appetite, instead of mysterious load.
- The crawled-but-never-referred list turns that spend into a content decision worth making.
- 01Upload this month's access logs from the Traffic page.
- 02Sort the daily rollups by crawl volume to name the heaviest operators.
- 03Take the per-operator crawl-to-refer ratio into the budget conversation.
Metrics this role tracks: Verified crawler hits · Crawl-to-refer ratio · Referral sessions
Growth lead attributing ChatGPT referralsThe sessions analytics files under Direct were your best new channel.
You report acquisition by channel and suspect AI-driven signups are hiding inside direct traffic.
- Seeded utm_source tags survive stripped referrers, so ChatGPT sessions stop reading as Direct.
- Six surfaces are attributed, ChatGPT, Perplexity, Gemini, Claude, Copilot, Meta AI, with a tagged-versus-organic split.
- Pair referral trends with mention rate and share of voice from scans to connect cause and effect.
- 01Paste the beacon snippet into your site from the Traffic page.
- 02Mint tagged URLs with the link builder and seed them into the AI surfaces you appear in.
- 03Watch the tagged-versus-organic split to see which attribution signal fired.
Metrics this role tracks: Referral sessions · Mention rate · Citation count
See referrals by sourceContent strategist closing the crawl-to-referral gapEngines read everything you publish and send readers to almost none of it.
You plan content against evidence and have watched AI crawlers sweep the site while referral numbers stay flat.
- The crawled-but-never-referred list names the pages engines consume without ever sending a reader back.
- Crawl-to-refer ratio read per intent separates a training crawler's normal appetite from an answer-index bot that cites nothing.
- Pairing the list with what engines say about you in scans separates a content problem from an answer problem.
- 01Upload last month's access logs from the Traffic page after proving domain ownership.
- 02Open the crawled-but-never-referred list on the Traffic page.
- 03Seed tagged URLs into your AI-facing links so future referrals attribute cleanly.
Metrics this role tracks: Crawl-to-refer ratio · Referral sessions · Verified crawler hits
Find the unreferred pagesOps engineer verifying bot traffic against published rangesVerify the wire, not the self-reported label.
You run WAF rules and firewall policy and want to know which bot traffic holds up against operator-published evidence.
- IP-range checks run against each operator's published file; reverse DNS counts only after forward confirmation.
- Operators publishing no verification mechanism stay marked unverified, the report says so rather than upgrading them.
- Nothing is excluded: totals stay comparable with raw logs while the confirmed share is reported beside them.
- 01Upload access logs from the Traffic page after proving domain ownership.
- 02Check the unverified rows against each operator's published method in the bot taxonomy.
- 03Compare dashboard totals against your raw logs, every hit is marked, none are dropped.
Metrics this role tracks: Verified crawler hits · Crawl-to-refer ratio · Referral sessions
Referrals
How do I track ChatGPT referral traffic?
Ask your analytics tool how many visitors ChatGPT sent and you will get an understated answer twice over. Most AI impact is zero-click, people read the brand inside an answer and act later without following the citation, and a meaningful share of the visits that do happen carry no referrer at all, because assistant apps are not browsers navigating from a page. Those sessions land in Direct and disappear into everyone else's traffic.
Capturing them takes two pieces working together. A first-party beacon on your pages records visits as they arrive, and a link builder mints tagged URLs whose utm_source names the surface you expect to read them. When a visit lands, the explicit tag wins over the referrer, deliberately, because the tag is the one signal that survives the trip through a desktop or mobile assistant. Six surfaces are classified by host or by tag: ChatGPT, Perplexity, Gemini, Claude, Microsoft Copilot and Meta AI.
Read the resulting number as a floor rather than a ceiling, and pair it with mention rate and share of voice from the monitoring side of the product. The tagged-versus-organic split tells you which attribution signal fired, useful context, because tags cover only the links you thought to seed, while everything else arrived by referrer host, which is precisely the signal that goes missing most often.
Crawlers
Which AI crawlers should I allow?
Eighteen crawlers from ten operators appear in the taxonomy. OpenAI, Anthropic, Perplexity, Google, Apple, xAI, ByteDance, Amazon, Common Crawl and Meta, and every hit is tagged with an intent: bulk crawling for training corpora, building an answer index, or fetching a page live because a user just asked a question. The intent changes how a crawl volume should be read.
A training crawler reading thousands of paths and referring nobody is behaving normally, it consumes far more than it sends back, by design. A burst from a live-fetch bot means somebody asked about you moments ago. And the bots are distinct controls rather than one blob: allowing the training crawler while blocking the search-index one is a coherent policy. Those calls are made in your own robots.txt, firewall or CDN, this surface observes and attributes; it never blocks on your behalf.
- Judge crawl-to-refer ratios per intent: a high ratio is normal for a training crawler and an anomaly for anything else.
- Answer-index crawlers such as OAI-SearchBot and PerplexityBot feed the indexes that answers cite, so their access and your citations are connected.
- Live-fetch hits, ChatGPT-User, Perplexity-User, Claude-User, are per-question events, so their timing maps to demand rather than to a schedule.
Verification
Can I trust a user-agent string?
A user-agent string is free text, so it is treated as a claim rather than a fact. Where an operator publishes a machine-readable IP-range file, the request IP is tested against those ranges. Where an operator documents reverse DNS instead, the IP is resolved, required to land under a documented hostname suffix, and then forward-confirmed back to the same address, a spoofer controls their own reverse-DNS record, so the lookup alone proves nothing.
Some operators publish no verification mechanism at all, and their hits stay marked unverified rather than being quietly upgraded to look decisive. Verification also never blocks ingestion: any failure records the row as unverified and the pipeline moves on. Totals therefore remain comparable with your raw logs, with the confirmed share reported beside the total, unconfirmed is not the same as fake, and deciding what to act on stays with you.
FAQ
AI Traffic questions, answered.
How do you tell a real AI crawler from something pretending to be one?
Where do the crawler numbers come from?
Some bots always show as unverified. Is that a bug?
Where do ChatGPT referrals come from when AI apps strip the referrer?
Does this show what engines say about my brand?
Why does uploading server logs require domain verification?
The beacon key sits in my page source. Is that safe?
Will you block AI bots for me?
Which plans include this?
How do I start collecting AI traffic data?
Keep exploring
Related features
The AI era of search
is already here
See what AI engines say about your brand before your competitors do. Start free today. No card required.

