Tracking your brand in AI answers means asking the questions your buyers ask — on ChatGPT, Gemini, Perplexity and Google's AI Overviews — and recording three things each time: whether you were named, whether you were recommended, and which sources the answer was built from. One check is a snapshot. A trend needs the same questions asked on a schedule.
Key takeaways
- Four surfaces carry most of the volume: ChatGPT, Google Gemini, Perplexity and Google AI Overviews. Each composes answers differently, so each is measured on its own.
- Three measures matter: mention rate (were you named), recommendation rate (were you put forward), and the split between sources you own and sources you do not.
- Answers vary between runs. A single check tells you little; the same set of questions asked daily or weekly is what shows movement.
- Start from 20–50 real buyer questions, not keywords. What a person types into ChatGPT is longer and more specific than a search query.
- A score only means something against competitors. Track the same questions for two or three rivals and read your share of the mentions.
Search used to be one box and ten links. A growing share of it is now a paragraph an AI writes for the person asking, and that paragraph either names your brand or it does not. Tracking is how you find out which, and whether it is getting better.
What "being tracked" actually means
There is no rank to check. An AI answer is prose, and your brand is either in it or absent. So tracking records three things for every question asked:
- Named. Did the answer mention your brand at all?
- Recommended. Was it put forward as an option, or only listed in passing — or discouraged?
- Sources. Which pages did the engine read to write the answer, and how many of them do you control?
The third is the one people skip, and it is the one that tells you what to do. If the answer was built entirely from review sites, forums and competitor pages, publishing more on your own site will not change it. That is a different piece of work, and you can only see it if you are looking at the citations.
The four surfaces
Most buyer questions run through four answer engines, and they do not agree with each other:
- ChatGPT — the largest, and the one most buyers mean when they say "I asked AI".
- Google AI Overviews — the block above the blue links. Reaches people who never intended to use an AI at all.
- Perplexity — smaller, but it shows its sources inline, which makes it the easiest place to see the citation picture.
- Google Gemini — its own answers, distinct from the Overview.
A brand can be strong on one and invisible on another, because each pulls from a different mix of sources. Track them separately or you average away the thing you need to see.
The three metrics
Mention rate. Of the questions you asked, how many named your brand at least once. This is your presence.
Recommendation rate. Of those same questions, how many put your brand forward as a choice. Being named in a list of eight is not the same as being the answer.
Share of voice. Of all the brands named across your questions, what fraction were you. This is the only figure that means anything, because it is relative: a 20% mention rate is bad in a two-horse race and good in a field of fifteen. To read it you have to track the same questions for your competitors, not just yourself.
Research on generative engine optimization has found that the levers which move these numbers most are citing primary sources, quoting named people, and including concrete statistics — not keyword density. Measuring first tells you whether you need any of that yet.
Start from questions, not keywords
A search query is two or three words. What someone types into ChatGPT is a sentence: "which tool should a small agency use to see if AI mentions its clients". Your panel of tracked questions should be written the way buyers actually ask — 20 to 50 of them, spread across the journey from "does a tool like this exist" to "how do I set it up".
If you have an existing keyword list, it is a starting point, but it is not the panel. The panel is the questions.
How often, and building a baseline
Ask the same panel on a schedule. Daily is the safe cadence, because the same question put to the same engine returns different answers on different days, and only repetition separates a real shift from that noise. Weekly works if the panel is large enough that one odd answer barely moves the average.
Your first run is the baseline: the mention rate, recommendation rate and share of voice you are starting from, per engine. Everything after is measured against it.
Doing it by hand, and where that stops
A spreadsheet works at small scale. Columns for the question, the engine, the date, named yes/no, recommended yes/no, and the sources you noticed. Ask, log, repeat.
It stops scaling at around 20 questions across four engines checked weekly — about 80 answers a week to read and record, at a fixed cadence, without missing a week. That is the point where a tool that does the collection for you starts to earn its cost, mainly because it keeps the schedule you would otherwise let slip.
Next
- How to check if ChatGPT recommends your brand — the manual method in detail
- How to see which sources an AI answer is built from — reading the citations
- Why ChatGPT doesn't mention your company — the usual reasons, in the order worth checking them
- Are AI visibility scores accurate? — how far to trust one number
- GEO vs SEO: what changes for a small agency — what carries over from search work
- How to write a GEO report for clients — turning the numbers into something to hand over
FAQ
- How do I check if my brand shows up in ChatGPT?
- Ask ChatGPT the questions a buyer would ask before choosing a product like yours, and note whether your brand is named, whether it is recommended, and what the answer cites. Repeat the same questions on other days — answers vary — and do the same for your main competitors so the result has something to compare against.
- How often should I track AI visibility?
- Often enough that a single unusual answer does not move your reading. Daily is the safe default because engines answer the same question differently from one day to the next, and only frequency averages that out. Weekly is workable if the panel of questions is large. Monthly is too sparse to tell a real change from noise.
- Which AI engines should I track?
- ChatGPT and Google's AI Overviews first, because they carry the most volume. Add Perplexity, which is where source citations are most visible, and Google Gemini. Microsoft Copilot and others matter less until you have the first four covered.
- Can I track AI visibility for free?
- Manually, yes. Keep a spreadsheet of your questions, ask them by hand, and log the outcomes. It stops scaling at around 20 questions across four engines checked weekly — that is roughly 80 answers to read and record every week, and it has to be done at the same cadence to stay comparable.
Sources
- GEO: Generative Engine Optimization (Aggarwal et al.) — arXiv, 2024
- GEO
- AI search
- measurement