Track the questions a buyer actually asks on the way to a decision, not keywords rewritten as sentences. For a single page, three to five prompts is enough. For a whole small site, twenty to thirty. Stop adding when new prompts stop surfacing new cited sources.
The number matters less than how often you run them. AI answers change from one run to the next, so twenty prompts run weekly tells you more than a hundred run once.
TL;DR
Three to five prompts per page, one to two per keyword phrase, twenty to thirty for a small site.
Six prompt types: category, problem, comparison, use case, recommendation, and negative.
Turn a keyword into a prompt by writing the question the person behind that keyword would ask out loud.
Repeat runs beat list length. If time is short, cut the list before cutting the repeats.
Record the answer text, which page was cited, and the sentiment, not just whether the brand appeared.
If nothing moves after six weeks, the problem is almost never the prompt list.
How to turn keywords into prompts
Short answer: take a phrase the page already ranks for and write the question a person would ask out loud to reach that page.
Start here: the phrases the page already ranks for. Each one is a question somebody typed, which is what makes it a usable prompt.
This is the part that decides whether the whole exercise is useful. Keywords rewritten as sentences produce prompts nobody asks. The phrase "cheap seo tools" becomes a real prompt as "what is the cheapest SEO tool for one website", because that is a question with a buyer behind it.
Where to get the raw phrases:
Search Console queries. Every row is a phrase a human typed. Filter for the question-shaped ones starting with what, how, which or best. There is no prompt volume data anywhere for AI assistants, so this is the strongest proxy that exists.
The keywords a page already ranks for. Same logic, and available even when Search Console is thin. Real people typed those phrases and landed on that page.
People Also Ask. The conversational layer on top of the same queries, already phrased as questions.
Customer emails and reviews. The last thirty messages before someone bought. Least polished wording, most honest.
How many prompts per page and per keyword
Short answer: three to five per page, one to two per keyword phrase, twenty to thirty for a small site.
Scope
Prompts
Why
One keyword phrase
1 to 2
One direct question, plus one comparison or use-case variant if the phrase is commercial
One page
3 to 5
The page's main question, one comparison, one use case, and one negative if the page sells something
A small site, five to ten pages
20 to 30
Enough to cover distinct buyer jobs without repeating the same intent
A catalogue with several audiences
30 to 60
Scale by audience and use case, not by page count
Count distinct buyer jobs, not phrasings. Ten ways of asking for the best tool in a category is one job, not ten. Most categories only have fifteen to thirty real jobs in them.
The stopping rule that beats a round number: keep adding until new prompts stop surfacing new cited sources. Two questions that pull the same five sources are one row for your purposes, because the fix for both is identical, which is to get named on those sources.
The six types worth tracking
Type
What it asks
Example
Category
What exists in this category
What are the best tools for tracking brand visibility
Problem
How do I solve this
How can I tell whether my brand appears in AI search
Comparison
Which of these two
Tool A vs Tool B for a small site
Use case
What fits my situation
Best SEO tool for a one-person business
Recommendation
What should I buy
What software should a small business use for AI search
Negative
What should I avoid
What is the worst tool for X
The first five follow the buying journey. The sixth is the one most lists skip and it earns its place: asking what is worst in a category shows whether the brand surfaces in an unflattering answer, and which sources the model reaches for when it needs criticism. Those sources are usually different from the ones it cites when recommending.
Category, comparison and negative prompts show how the models see the brand. Use case and recommendation prompts show whether the brand survives to the shortlist.
Why runs matter more than the length of the list
Short answer: the same prompt asked five times returns different answers, so one pass is a screenshot rather than a measurement.
These models are not deterministic. Ask the same question twice and the brands named, their order, and the citations can all change. A number from one run cannot be compared to a number from another run without knowing the spread.
Twenty prompts run ten times says more than two hundred run once, because only the first can separate a real position from run-to-run noise. If time or budget is short, cut the prompt list before cutting the repeats.
This has an awkward consequence worth stating plainly: most tools in this category sell prompts rather than runs. The plan caps how many questions you can track, while the thing that decides whether the data means anything is how often each one is asked.
What to record for each prompt
One row per prompt, one column per engine. Mentioned and cited are recorded separately, because they move independently.
Short answer: the answer text, the page that was cited, and the sentiment. Not just whether the brand appeared.
A yes-or-no record cannot tell you why something changed. The minimum row that survives contact with reality:
The date, and the prompt exactly as it was asked.
The engine and the model name, and whether web search was on.
Mentioned: whether the brand was named in the text.
Cited: whether a page was linked, and which page.
Sentiment: whether the mention was positive, neutral or negative.
The other brands named in the same answer.
The full answer text.
Mentioned and cited are separate events and they move independently. A brand can be named while the link underneath points at a review site, which means the answer does the recommending and someone else collects the visit. Counting them as one number hides the most useful thing on the page.
Sentiment matters for the same reason. Appearing in an answer about what to avoid is still an appearance, and a tracker that only counts mentions will report it as progress.
The answer text is what turns an unexplained drop into a diagnosis. Without it, a real loss and a random fluctuation look identical on a chart.
What to do when nothing moves
Short answer: after six weeks with no change, the problem is almost never the prompt list.
Work through these in order, because the cheapest checks rule out the most common causes.
1
Rule out variance first
Compare the last three runs, not the last two. If this week's number lands inside the range of the previous few, nothing actually moved.
2
Check whether you appear anywhere at all
Run a free brand check before anything else. If the brand returns zero everywhere, the issue is not the prompts, it is that nobody mentions the brand yet, and no amount of tracking changes that.
3
Check that the page can be read
If the text only appears after scripts run, or a crawler is blocked, the answer engines never saw it.
4
Look at who is cited instead of you
The same five to ten domains usually appear across every prompt in a category. That list is the actual work: getting named on those pages moves the answer far more than publishing another article.
5
Check the prompt is one buyers ask
If a prompt was invented rather than taken from real query data, a zero on it means nothing.
6
Give it thirty to forty-five days
First movement on narrow questions usually appears in that window. Two weeks is not a measurement.
If the brand is cited but not mentioned, the page is being used as a source without the model naming you, and the fix is on the page: state the brand plainly next to the fact worth quoting. If the brand is mentioned but not cited, the fix is also on the page: make the answer easy to lift and easy to attribute.
Doing this without an hour of copying and pasting
The answer itself, kept per engine: date, model, web-search flag, status and the sentence naming the brand.
Running twenty prompts across three engines by hand, every week, is about an hour of work, and it is the part people stop doing after a month.
IvaBot runs the prompt list across ChatGPT, Perplexity and Google AI and records, for each prompt and each engine, whether the brand was mentioned, whether a page was cited and which page, the sentiment of the mention, the competitors named in the same answer, and the full answer text so you can read what the model actually said. The first prompts are derived from the keywords the page already ranks for, which is the Search Console logic above applied automatically, and the rest are yours to add. Fifty prompts a month costs about five dollars, and new accounts get three free credits with no card.
Disclosure: IvaBot is my own tool.
FAQ
Which AI prompts to track for SEO?
The questions a buyer asks on the way to a decision, drawn from Search Console queries and the keywords a page already ranks for, split across the six types above.
How do you actually keep track of prompts that work?
Freeze the list and rerun the same prompts on a schedule, recording the answer text each time. A prompt works when it consistently returns your brand across runs, and that can only be seen across several runs of the same wording.
How to track ChatGPT prompts?
Ask the same fixed set in ChatGPT with search enabled, record whether the brand is named and which sources are cited, and repeat weekly. A ChatGPT prompt tracker automates the repetition and keeps the answer text so the runs can be compared.
How many prompts per page?
Three to five. The page's main question, one comparison, one use case, and one negative if the page sells something.
Is there prompt volume data?
No public equivalent of keyword volume exists for AI assistants. Search Console query data is the closest proxy, which is why prompts are best derived from it.
Should prompts be full sentences or keyword-style?
Track a mix. People using voice or a phone tend to write full sentences, people typing quickly use short phrases.
How often should I run them?
Weekly is enough for most small sites. Monthly is too sparse to separate a real change from noise.
Write down five questions a customer would ask before they know your name, run them this week, and record what came back. That is the whole first pass.
IvaBot runs the list across ChatGPT, Perplexity and Google AI and keeps the answer text, so the next run can be compared with this one. Three free credits, no card. Check your page.