A tracked prompt is a question you have decided to ask AI engines on every scan, stored against a project. It is the smallest unit of measurement in the product: every mention rate, position, citation and share-of-voice figure is derived from answers to prompts, so the prompt set defines what your data can possibly tell you.
This article covers the mechanics — what happens when you activate a prompt, how the limits work, what order they run in — and then the selection judgment.
What activating a prompt actually changes
By default, a scan builds its own discovery prompts from the project’s brand name, industry, competitor list and region. They look roughly like this:
- “What are the best [industry] tools available right now? List the top options.”
- “I’m evaluating [industry] solutions. Which would you recommend and why?”
- “Is [Brand] a good [industry] option? How does it compare to alternatives?”
- “What are the best alternatives to [first competitor] for [industry]?”
Those are reasonable starting questions and a poor long-term measurement instrument, because they are generic by construction.
The moment a project has one or more active tracked prompts, the scan runs those instead — the auto-generated set is not used at all. This is a replacement, not an addition. A project with a single active prompt scans exactly that one question and nothing else, which is a common and avoidable way to end up with a very narrow first week of data.
Deactivating every prompt returns the project to auto-generated discovery prompts.
The two limits, and why they are different
There are two caps and confusing them causes real problems.
Active tracked prompts bounds how many prompts a project may keep active — what you store. Prompts per scan bounds how many one scan will actually run. The second is a runtime guard first and a plan perk second: prompts run sequentially per engine, and a long enough set makes a scan exceed the worker’s time budget.
At the time of writing:
| Plan | Domains | Tracked prompts | Credits / month | Refresh |
|---|---|---|---|---|
| free | 1 | 10 | 10 | manual, monthly |
| team | 5 | 50 | 250 | weekly |
| business | 20 | 200 | 1,000 | weekly |
| agency | 100 | 1,000 | 5,000 | weekly |
Two rules generate that whole table: every domain gets 10 tracked prompts, and every prompt is funded for 5 refreshes a month — enough for the weekly schedule with room for scans you run yourself.
One credit is one prompt, checked once, across every engine you have enabled. Not one per engine: asking “best CRM software” on all seven costs one credit, not seven. So your monthly consumption is prompts × refreshes, and turning engines off does not buy you more of them — it just gives you less data for the same price.
A ceiling of 25 prompts per scan applies per project on top of the plan figure, and prompt text is capped at 500 characters. Check the pricing page for the current numbers.
The relationship between the two columns is the part to internalise. On a business plan you may keep 50 prompts active, but any single scan runs 15 of them. The active cap exists precisely so that gap stays honest: without it you could activate hundreds of prompts, watch every scan silently cover a fraction, and have nothing in the interface tell you so.
Which prompts a scan picks when you have more than it runs
Active prompts are loaded oldest-first, by creation time, and the scan takes the first N. The prompts you created earliest are the ones that get scanned every time.
That is a stable, predictable rule, and it has one practical consequence: if you are near your per-scan limit, adding a new prompt does not displace an old one — it queues behind it. When a prompt stops earning its place, deactivate it rather than leaving it active and hoping the newer one wins. Prompt hygiene is not optional once your active set exceeds your per-scan set.
Budgeting: queries are engines × prompts
One scan costs one query per engine per prompt. Five prompts across six engines is thirty queries, not five. This is the arithmetic that decides how many prompts you can afford at your monitoring cadence, and it is worth doing before you build the set rather than after.
Answers that fail — a provider timeout, an engine error — are excluded from the charge and from the free-tier quota. You are billed for delivered answers, not attempted ones.
Engine selection is a real budget lever alongside prompt count. Multi-engine monitoring covers how much answers diverge between engines, and which AI engine matters most is the argument for narrowing when budget is tight.
Choosing the questions
The selection framework is the same one that governs any monitoring set, and building your prompt monitoring strategy covers the full taxonomy and allocation percentages. The version tuned to a per-scan limit in single digits:
Write questions, not keywords. Engines are asked things like “What is the best CRM for a five-person sales team?” Search boxes get “crm small business”. Google’s own guide to optimizing for generative AI features reflects the same shift — there is no separate keyword list for AI surfaces, and its companion documentation on AI features and your website states there are no additional requirements or special optimizations to appear in them. What changes is the phrasing of the question, and your prompts should match how buyers actually phrase it.
Weight unbranded over branded. Branded prompts (“Is [Brand] any good?”) almost always mention you — that is what makes them nearly useless as a visibility signal and genuinely useful as an accuracy signal. Unbranded category questions are where competitive position is actually decided. See branded vs unbranded AI queries. With a limit of ten prompts, spending more than two on branded questions is usually a mistake.
Include the comparison you lose. “Alternatives to [market leader]” is the highest-signal prompt most brands do not track, because it is the question a buyer asks at exactly the moment a vendor shortlist forms.
Cover one intent per prompt. A prompt bundling two questions produces an answer that is hard to score and harder to compare over time.
Do not paraphrase for coverage at small set sizes. Three phrasings of the same question is a legitimate technique for probing semantic coverage, but it consumes three of your ten slots. Below about twenty active prompts, breadth of intent beats depth of phrasing.
Let your Search Console data pick some of them
If you have connected Search Console, you have a list of questions your market demonstrably asks, with volume attached. The Search↔AI gap report joins those queries to your tracked prompts by token containment, and a query that matches no prompt is counted as unmatched — no evidence either way, because nobody asked the engines about it.
A large unmatched-query count is the single best prompt-selection input available, and it is empirical rather than a guess. Reading your Search↔AI gap report explains how the matching works and why an unmatched query is never reported as a finding; why rankings and citations disagree covers why the two surfaces need separate measurement in the first place.
Prompts, regions and competitors
A project’s region instructs the engines to answer for that market and biases the auto-generated prompts, so you generally do not need to write “in Germany” into every tracked prompt — see geographic variation. Separate markets are usually better served by separate projects, since that also separates the scoring.
Competitor benchmark scans deliberately run the same prompt universe as your own brand’s scans. That is what makes the comparison score-against-score on identical questions rather than two differently-worded sets producing an incomparable number. Competitor benchmarking and the competitor benchmarking guide cover how to read the result.
A workable starting set
For a first project on a plan that runs ten prompts per scan:
- Four category-intent questions — the “best X for Y” phrasing your buyers use, varying the qualifier (company size, industry, constraint) rather than the wording.
- Two alternatives-to-competitor questions, naming your two most frequently encountered rivals.
- Two problem-first questions describing the pain rather than the category.
- One branded accuracy question — “What does [Brand] do and what does it cost?” — as your hallucination tripwire.
- One question drawn from your highest-impression unmatched Search Console query.
Run that unchanged for a month before editing it. Answers move between runs on their own, and a set you keep rewriting produces a trend line you cannot read.
Frequently Asked Questions
Do tracked prompts add to the automatic prompts or replace them?
They replace them entirely. As soon as a project has at least one active tracked prompt, scans run only tracked prompts and the auto-generated discovery set is not used. Deactivating all of them restores the automatic behaviour. This is why activating a single prompt narrows a project’s data rather than enriching it.
What happens if I activate more prompts than my plan runs per scan?
The scan runs the oldest-created active prompts up to the per-scan limit and skips the rest — every scan, consistently. The active-prompt cap exists so that this truncation stays visible rather than letting you accumulate a long active list that only ever partly runs. If a prompt no longer earns its slot, deactivate it.
How many prompts should I start with?
Fill your per-scan limit and no more — five on free, ten on team. A set you run every scan for a month produces a readable trend; a larger set that only partly runs each time does not. Expand once you can point at a specific question your current set cannot answer.
Should I track branded prompts?
One or two, as an accuracy check rather than a visibility metric. Engines nearly always mention your brand when asked about it by name, so branded prompts tell you little about competitive position — but they are the fastest way to catch an engine stating wrong pricing or a wrong category for you.
