The manual method, run on my own site with a result of zero: which prompts to use, why one run proves nothing, what to record, and how to calculate the number.
Try IvaBot for your site
Find what's hurting your SEO and what to fix first. Free to start, no credit card.
To measure brand visibility in ChatGPT, pick a fixed set of non-branded prompts your buyers would actually type, run each one at least three times in a clean session, and record two separate events for every run: whether the brand was named, and whether your site was cited as a source. Visibility is the share of runs where the brand appeared, counted against the total number of runs rather than the number of prompts.
Everything below is that method in detail, run on my own site, where the honest answer came back as zero. The screenshots are from those runs, including the one where I got the measurement wrong.
TL;DR
Measure non-branded prompts. Asking ChatGPT about your own brand tests nothing worth knowing.
Use a temporary chat with memory off. An account that has researched your category will name your brand back to you.
Three runs per prompt is the minimum. A single run carries no information.
Record mentions and citations as two separate columns. They move independently.
Divide by the number of runs, not the number of prompts.
Count each engine separately. Averaging hides the finding.
A zero is diagnostic. The list of domains cited instead of you is the actual task list.
What brand visibility in ChatGPT means
Short answer: the share of relevant AI answers where your brand appears, measured across repeated runs.
There is no position to hold. ChatGPT either includes a brand in an answer or leaves it out, so the unit of measurement is presence rather than rank.
Two different events hide inside that presence, and they move independently. A brand mention is the model naming the brand in the text of the answer. A citation is the model attaching your URL as a source. A brand can be named in every answer and cited in none, which usually means the model learned about it from other people's pages rather than from yours. The reverse also happens, where a page is cited as a source for a fact while the brand behind it is never named. Counting these two as one number is the most common reason two people measuring the same brand in the same week get different results.
Sentiment sits on top of both. Being named as the cheap option and being named as the category leader are the same mention in a counter and two different outcomes for the business, so the context of the mention gets recorded alongside the fact of it.
Why your usual metrics do not answer this
Short answer: rankings do not carry over, and analytics only sees the conversations that ended in a click.
One figure passed around in this space puts the chance of a first-place Google result appearing in the ChatGPT answer to the same question at roughly one in five. Treat that as an observation rather than a law, since it comes from a practitioner comment rather than a published study and I have not been able to trace it to a primary source. What holds up in my own testing is the direction: the model assembles an answer from what it trusts about a category rather than reading the results page in order, so a strong ranking predicts less than it feels like it should.
Analytics does not close the gap either. GA4 sees the click when someone follows a link out of ChatGPT, so it counts the traffic that survived the answer. It cannot see the far larger number of conversations where the brand was named, the reader decided, and no click happened at all.
A note on the numbers in this guide. This is one person's method, tested on my own site and on client work, combined with published studies where they exist and with what practitioners report in public threads. Where a figure is mine, I say so. Where it is someone else's, I say whose and how solid it looks.
Step 1: Build the prompt list
Short answer: fifteen to thirty non-branded prompts, phrased the way a buyer would ask before they know your name.
Measure non-branded prompts. Asking ChatGPT about your company by name tests whether the model has heard of you, which is a question with an obvious answer and no strategic value. The real test is the question your buyer types before they know your name, in the form of best tool for a job, alternatives to a competitor, or a problem stated plainly.
Fifteen to thirty prompts is enough for one small site. Fewer than that and one odd answer moves the whole number. More than that and the manual method collapses under its own weight, for reasons the run count below makes clear.
The five prompts used throughout this guide were these:
What SEO tool lets me pay only for what I use?
How can I audit my site's SEO without a subscription?
What are the cheapest AI visibility tracking tools for a single website?
How can I track brand mentions in AI search engines?
What tool tracks both Google positions and AI citations?
The selection method matters more than the count, and it is written out separately in the guide on deciding which AI prompts to track, including the six buying-stage types and the negative prompts most lists leave out.
Step 2: Clean the session
Short answer: a temporary chat with memory switched off, or the measurement reports your own history back to you.
Turn off ChatGPT memory before the first prompt and use a temporary chat for every run. An account that has been used to research your own category is an account the model has already learned from, and it will name your brand back to you because you taught it to. Perplexity has no persistent memory, so a separate session is enough there.
Memory off in Personalization. Do this once, then use a temporary chat for every run after it.
Here is what a contaminated run looks like, from my own first attempt. I asked which SEO tool lets you pay only for what you use, in my normal logged-in account. The answer recommended a tool with the line that for my IvaBot project specifically it would pick that one, and put IvaBot in a comparison table under best for. I never typed the word IvaBot.
A run in a normal logged-in account. The prompt named no brand. The model supplied one from account memory.
Read as a measurement, that answer says the brand is visible. It says nothing of the kind. It says the account owner had discussed the brand before. Every later run in the same account carried the same contamination, which is why those runs agreed with each other and disagreed with the clean ones. Contaminated sessions look stable, and looking stable is exactly what makes them convincing.
Region and language change the answer as well, so fix both before the first run and keep them identical afterwards. A measurement taken from a different country is a different measurement.
Step 3: Run each prompt three times
Short answer: three separate temporary chats per prompt, because one run of one prompt carries no information.
How visibility is measured when the same question returns a different answer each time is the objection that stops most people, and the answer is repetition. Three runs per prompt per engine is the working minimum, and five is better where the answers disagree. Open a new temporary chat for each run rather than repeating the prompt in the same window, because inside one chat the model can see its own previous answer and will echo it.
Run one. The Temporary chat label in the input field is the only reliable confirmation that memory is out of the picture.Run two. Same prompt, minutes apart, different brands.Run three. The overlap between the three lists is partial, which is the whole reason a single check proves nothing.
If you had run this once and stopped, you would have recorded whichever set of brands happened to appear that time and treated it as the state of your category. Three runs make the variance visible, and the variance is the finding.
Step 4: Run it with search on and off
Short answer: two runs of the same prompt, because training data and live retrieval answer different questions.
With retrieval switched off you are reading what the model absorbed during training. With it on you are reading what it found on the web this minute. Both are real. They are not the same measurement, and mixing them in one column is how a number stops meaning anything.
Search off. Brands are named from training data with no sources attached, so every appearance here is a mention and none is a citation.Search on. Source chips appear next to the brands, which is what turns a mention into a citation and gives you a domain to record.
In a test across 57 prompts, the recommendation share came out at 0.402 with forced retrieval against 0.256 without it, a gap of 36 percent, and the retrieval-backed numbers were about three times less stable week to week. That instability is the price of freshness, and it is another argument for repeated runs.
Step 5: Record the same fields
Short answer: one row per run, with mentions and citations in separate columns.
Use one row per run, not per prompt. Collapsing three runs into one row destroys the variance you just paid for.
Field
What goes in it
Date
The date of the run
Prompt
The prompt exactly as typed
Engine
ChatGPT, Perplexity or Gemini
Web search
On or off
Mentioned
Yes or no, brand named in the text
Cited
Yes or no, your URL attached as a source
Sentiment
Positive, neutral or negative
Position in answer
First, inside the list, or in passing
Competitors named
Every other brand in the answer
Sources cited
Every domain the answer linked to
The last two rows are the ones people skip and later wish they had kept. Competitors named tells you who owns the category in the model's view. Sources cited tells you which pages you would have to appear on to change that, and it is the only column that converts a score into a task.
Step 6: Calculate the number
Short answer: mentions divided by total runs, with citation rate kept as its own figure.
Calculating brand visibility is a division, and the denominator is the part that gets fudged.
Visibility rate is the number of runs where the brand was mentioned divided by the total number of runs. Fifteen prompts, three runs each, one engine gives 45 runs. Named in nine of them is a visibility rate of 20 percent.
Citation rate is the number of runs where your URL appeared as a source, divided by the same total. Keep it separate rather than folding it into the first number.
Share of voice is your mentions divided by the mentions of every brand counted in the same answers. It is the number most tools lead with and the number that means least on its own, because half of a tiny category is worth less than a tenth of a large one.
Count each engine separately before averaging anything. ChatGPT and Perplexity draw on different sources, so a brand can hold a solid rate in one and zero in the other, and an average hides exactly the finding you needed.
What the number looked like on my own site
Short answer: zero, and the useful part was the list of domains cited instead.
I ran this on ivabot.xyz in August 2026. Fifteen non-branded prompts, written to match the category rather than the brand. Across those prompts the brand was named zero times, the site was cited zero times, and share of voice came out at zero percent. A separate manual check in ChatGPT on the cheapest tool for a single website returned the same nothing, which matters, because it means the zero is real rather than a tracking artifact.
For scale, in the same measurement Semrush took 32 mentions, 53 percent brand coverage and 86 percent share of voice, and Otterly took 5 mentions and 14 percent share of voice.
The useful part was not the zero. It was the source column. The domains cited when the models answered those prompts were onelittleweb.com, youtube.com and vezadigital.com at about 15 percent citation share each, then ranklytics.ai, gomega.ai, behindrankings.com, cited.so, freddiechatt.com, developers.google.com and reddit.com. Those are named pages with named authors, and appearing in them is a task with a deadline. A visibility score is not.
Three free signals worth adding
Short answer: Search Console for Google surfaces, Bing Webmaster Tools for Copilot, and branded search for humans.
Google Search Console reports how often your pages appeared in Google's AI features. It is first-party data about your own site and it costs nothing. Three limits before reading anything into it: it reports impressions only, with no clicks and no queries, so you see that a page appeared without knowing what was asked; it covers Google surfaces and nothing else, so ChatGPT and Perplexity stay invisible in it; and it can only surface pages Google already retrieves. On my own site it reported two appearances across three months, which is consistent with the zero measured through prompts rather than a contradiction of it.
Three months of appearances in Google's AI features. A number this small is still information, because it confirms the pages are reachable and simply not chosen.
Bing Webmaster Tools is the second one. It reports citations for grounding queries, showing which of your pages were cited and how often across the Microsoft and Copilot surfaces, which use ChatGPT as their core. It does not cover the ChatGPT app itself and gives no totals to compare against, so the value is in the month-over-month direction rather than the absolute figure. It is worth setting up precisely because it can show citations for a site that measures zero in every prompt-based check.
Citations reported in Bing Webmaster Tools. First-party, free, and available even where the prompt-based number is zero.
Branded search volume is the third. If the work is landing, more people type your brand name into Google after hearing it in an answer, and that lift shows up in Search Console before revenue does.
Together these three cover the Google side, the Microsoft side and the human side. The prompt log is what covers ChatGPT itself, which is the one surface none of them reach.
What to do when the number is zero
Short answer: work off your own site, because mentions elsewhere predict inclusion far better than links do.
A zero means the model has nothing to work with, and the fix is mostly off your own pages.
Ahrefs studied 75,000 brands and found that branded web mentions correlate with visibility in AI Overviews at 0.664, while backlinks correlate at 0.218, roughly three times weaker. Brand anchors came in at 0.527 and branded search volume at 0.392, so the three strongest signals in the study were all off-site. A follow-up in December 2025 extended the work to ChatGPT and Google AI Mode and found mentions on YouTube correlating strongest of all, at 0.737.
Read those figures with the caveats the study itself carries. Correlation is not causation, and large brands naturally have both more mentions and more visibility, so brand strength may be driving both. The sample covered brands with a Domain Rating above 40, which are already established. The first study measured Google AI Overviews rather than ChatGPT. What the numbers support is a direction, and the direction is that being talked about elsewhere predicts inclusion far better than links into your own site.
In practice that means three things. Get named on the pages your source column already identified. Find the mentions you already have that carry no link, which is covered in the guide on unlinked brand mentions. Make the pages you own easy to quote and easy to attribute, which is covered in the guide on getting mentioned in Google's AI search.
Where the manual method breaks
Short answer: at the point where repeating it weekly costs more than the answer is worth, and where old answers are gone.
Manual measurement is the right way to start and the wrong way to continue, and it fails in a specific place rather than all at once.
Fifteen prompts, three runs, three engines is 135 conversations for one measurement. Weekly, that is 540 conversations a month before anything is written down. Fifty prompts tracked daily across three engines is 4,500 model calls a month, which is where a spreadsheet stops being a tool and becomes a job.
The second failure is quieter. A number you cannot reopen is a number you cannot defend. When the rate drops from 20 percent to 12 percent, the only question that matters is what the answers actually said last time, and a spreadsheet cell reading no does not tell you. Storing the full text of every answer, with its date, engine and web search flag, is what turns a drop into a diagnosis.
Below is the same measurement in my own tool, using the identical five prompts from Step 1. The manual version took nine conversations and about half an hour. This version is one run, and the answers stay readable afterwards.
The same five prompts, entered once instead of typed into nine separate chats.Each answer is kept in full with its date and engine, so next month's number can be compared against the text that produced this month's.
What tool is best for measuring brand visibility in ChatGPT?
Short answer: it depends on how often you need the answer, not on which dashboard looks best.
Disclosure before the table: IvaBot is my own tool, and it is marked as such in the table like everything else. Prices are summer 2026.
Feature
Semrush
Ahrefs
Otterly AI
IvaBot (mine)
Covers ChatGPT
Yes
Yes
Yes
Yes
Other engines
Perplexity, Gemini, AI Overviews, AI Mode
AI Overviews, AI Mode, Perplexity, Copilot, Gemini
Three more included, two as paid add-ons
Perplexity, Gemini
Prompts included
50 on Semrush One Starter
Prompt volume rather than a set list
15 on the entry plan, 100 on the next
10 prompts per credit
How often it runs
Daily
Monthly data refresh
Daily
When you run it
Full answer text kept
No
No
No
Yes
Sentiment per mention
Yes
No
Yes
Yes
Competitors in the answer
Yes
Yes
Yes
Yes
Google positions and backlinks
Yes
Yes
No
Yes
Price
$199 and $299 a month for Semrush One, around $99 for the toolkit alone
Brand Radar around $199 per engine, around $699 for all
$29 a month for 15 prompts, $189 for 100
$5 for 5 credits, $34 for 60, paid when you need it
Free entry
Free checker, then a card
Free checker, no signup
Full report, no card
3 credits, no card
Read it against how often you need the answer. If you report to a client every Monday, a daily tracker earns its subscription, and the choice is between Semrush and Otterly depending on whether you also need Google data. If you check your own site once a month, thirty prompts through IvaBot costs about three dollars for that month and nothing for the months you skip, which is the case the subscription tools price badly. If you run an agency across many clients, none of the pay-per-use options fit, and the honest answer is a seat-based tool.
The trade-off on IvaBot is stated plainly. There is no daily automatic run, so the trend line has points where you ran it rather than points on a schedule, and there are no alerts. What it does have is the stored answer behind every data point, which is the thing a spreadsheet cannot keep and most dashboards throw away. Three free credits, no card.
FAQ
What does brand visibility mean in AI search?
It is the share of AI answers to relevant questions where your brand appears, measured across repeated runs rather than a single check. It replaces ranking as the unit of measurement, because AI answers have no positions to occupy.
How is brand awareness measured differently from AI visibility?
Brand awareness is measured against people, through surveys, branded search volume and direct traffic. AI visibility is measured against models, by running prompts and counting appearances. They correlate, since a brand people discuss is a brand models have read about, but they are separate numbers with separate methods.
How many prompts do I need to measure AI visibility?
Fifteen to thirty for a small site, run three to five times each. The run count matters more than the prompt count, because a single run of any prompt is noise.
Why does ChatGPT give a different answer every time I ask?
The model samples its output rather than looking up a fixed result, and web retrieval adds a second source of variance by pulling different pages on different days. Repetition is the only way to separate a real change from normal variance.
Can I track ChatGPT visibility for free?
Manually, yes, and the method above costs nothing but time. Search Console shows appearances in Google's AI features and Bing Webmaster Tools adds citation data for the Copilot surfaces, both first-party and free. Neither covers the ChatGPT app, which is why the prompt log still does the main work. Free tool tiers exist as well, and four of them were tested without a card in a separate guide.
Does ranking first on Google mean ChatGPT will mention me?
No. The figure often quoted is about a one in five chance, and even taken loosely the direction is clear: the model weighs what other sources say about a brand more heavily than where a page sits in the results.
Should I measure with web search on or off?
Both, at least once. With it off you measure what the model already believes about your category. With it on you measure what it can find today. The two numbers differ by a wide margin and answer different questions.
What is a good AI visibility score?
There is no benchmark that travels across categories, because share of voice depends on how many brands the models consider relevant. The useful comparison is your own number against the previous run, measured the same way.
Write down five questions a buyer would ask before they know your name. Turn off memory, open a temporary chat, and run the first one three times. Record what came back. That is the whole first pass, and it takes about fifteen minutes.
IvaBot runs the list across ChatGPT, Perplexity and Google AI and keeps the full answer text, so next month's number can be compared with this one. Three free credits, no card. Check your site.