An AI visibility score is directional, not exact. The same question returns different answers on different days, and one study found less than a 1 in 100 chance of getting the same list of brands twice. A score is trustworthy when it is a rate across many questions, repeated over time, and it is not when it reports one run's position as fact.
Key takeaways
- SparkToro's study (2,961 runs, 12 prompts, ChatGPT, Claude and Google AI) found less than a 1 in 100 chance of an identical brand list and less than 1 in 1,000 of an identical order.
- That makes a rank position in an AI answer a poor number to report. A visibility rate across many questions, repeated, is the more defensible one.
- A single run is a snapshot. It can be useful for showing a client a gap. It cannot show that anything changed.
- Say what the number is: the share of answers that named the brand, on this date, on these questions and engines.
- Movement across several runs is the evidence. One run's change is noise until the next run agrees.
Ask ChatGPT for the best tool in your category. Now ask again in a fresh chat. Compare the two lists.
They are rarely the same. Different brands appear, and the same brands show up in a different order. That one observation is the whole accuracy problem with AI visibility scores, and it is worth understanding before you hand a client a number.
What the study found
In early 2026 Rand Fishkin at SparkToro published a study on this. It covered 12 prompts across consumer and business categories, run 2,961 times in total by 600 volunteers, on ChatGPT, Claude and Google's AI Overview and AI Mode, between November and December 2025.
Two results stand out. There was less than a 1 in 100 chance that ChatGPT or Google's AI would return an identical list of brands across repeated runs. And there was less than a 1 in 1,000 chance of returning the same list in the same order.
Fishkin's advice follows from that. Do not report a ranking position in an AI answer. Report a visibility percentage, measured across dozens to hundreds of prompts, each run more than once. It is a small study of 12 prompts, and its author is a vendor of audience research tools, so treat the exact figures as one credible data point and not a law. The direction has held up in every tool we have looked at.
What a score is, then
A visibility score is a rate: of the answers collected, the share that named the brand. It is an estimate, and like any estimate it has a margin.
Two things shrink the margin. More questions, because one unusual answer matters less among fifty than among five. And more runs, because repeated collection averages out the sampling.
Two things do not. A tool that reports "you rank third in ChatGPT" from a single answer. And a tool that reports a rise or a fall from one run to the next without saying that the difference may be within the noise.
What to tell a client
Use words that match what you have.
On a first run, say it is a snapshot. "On 12 September, across these 40 buyer questions on four engines, your brand was named in 11 answers." That is a fact about a date. It is enough to show a gap, and it is fine for a pitch, since running it on a prospect is the whole point.
On the second run, wait. If the rate moved, say it moved and that you will confirm on the next. If it did not, say it held.
On the third, you can talk about a trend, if the runs agree. That is the moment a number becomes evidence, and it is why a re-run schedule matters more than a clever first run.
Do not promise the reason. The number tells you where a brand stands, and the sources behind each answer tell you what might change it. No one has shown that a particular fix will move a particular answer.
Where GeoShark stands
GeoShark reports are directional, and they say so. A single run is a snapshot; answers vary between runs. The value is in the sources it exposes and in the movement over time, which is why a run is repeated on a schedule and shown against the last one. If a tool tells you its score is exact, ask how it handled variance.
For the mechanics of checking by hand, see how to check if ChatGPT recommends your brand. For the report itself, see how to write a GEO report for clients, and for choosing a tool, how to choose an AI visibility tool for agencies. New to the term? Start with what GeoShark and GEO are.
FAQ
- Why do AI answers change between runs?
- The model samples its reply instead of producing one fixed output, and engines that search the web retrieve different pages at different times. Both effects show up in the brand list. SparkToro found the chance of two runs giving an identical list of brands was under 1 in 100.
- How many times should a prompt be run to get a reliable number?
- There is no agreed figure. SparkToro's advice is to measure visibility as a percentage across dozens to hundreds of prompts run multiple times, not to rely on any single answer. The more questions and the more repeats, the smaller the effect of one odd answer.
- Can I tell a client their AI visibility went up?
- Only when more than one run agrees. A rise in one run can be sampling noise. If the rate is higher across two or three consecutive runs on the same questions, you can say it moved. Say the size of the panel as well, since a rise from 2 answers to 3 is not the same as 20 to 30.
- Is a lower-confidence score still worth showing a client?
- Yes, if you label it. A snapshot that shows a brand named in one of ten questions is enough to start the conversation. Say that it is a snapshot, give the date and the question set, and promise a re-run rather than a verdict.
Sources
- AIs are highly inconsistent when recommending brands or products — SparkToro (Rand Fishkin), 2026
- GEO: Generative Engine Optimization (Aggarwal et al.) — arXiv, 2024
- AI visibility
- measurement
- accuracy