Guides

How to choose an AI visibility tool for agencies

Written by

GeoShark

Last updated: 21 September 20264 min readReviewed by GeoShark Team

Share

Choose an AI visibility tool by six things: what it meters, which engines it reads, how it handles answers that change between runs, whether it shows the sources behind each answer, what you can hand a client, and whether it writes any of the work. Price per client depends on the first one. Test two tools on one client for two weeks.

Key takeaways

  • Work out cost per client, not headline price. A tool that charges per prompt and one that charges per brand give different answers as your roster grows.
  • Ask which engines it reads and how it gets the answer. An answer from the live product is not the same as an answer from an API.
  • AI answers vary between runs. SparkToro's study found less than a 1 in 100 chance of the same brand list twice, so a tool that reports a single run as fact is overstating.
  • The score tells a client where they stand. The sources tell you what to do. Prefer a tool that shows both.
  • Ask what you can hand over on the last day of the month, and whether it costs extra to put your own branding on it.

Plenty of tools now sell some version of "see whether ChatGPT recommends your client". They look alike on a landing page. They are not alike when you have eight clients and a client asks "so what do we do?" on a Thursday.

This is a buyer's checklist, not a ranking. It names no rivals and quotes no rival prices, because those change monthly and a stale table is worse than none. It gives you the six questions that separate one tool from another once you are past the demo.

1. What does it meter?

Some tools charge per brand, some per seat, some per tracked question. None is wrong. Each behaves differently as you add clients.

Do the sum before you look at the price list. Decide how many buyer questions you would track for a typical client, 40 to 60 is a workable panel, multiply by the number of clients you have or expect, and price that. A per-brand tool can be cheap for one large client and expensive for six small ones. A per-prompt tool is the reverse. GeoShark meters on prompts, with plan tiers of 100, 300 and 800. Whichever tool you look at, the point of the sum is that you do it.

2. Which engines, and how does it get the answers?

Ask for the list, and ask whether each engine is read the same way. An answer copied from the live product is what your client's buyer sees. An answer from a developer API can differ from it.

Coverage to expect: ChatGPT, Google AI Overviews, Perplexity and Gemini. Anything beyond that is a bonus until those four are read reliably.

3. What does it do about variance?

This is the question a demo skips. Ask a tool the same question twice and you get two answers. SparkToro ran 2,961 prompt runs across ChatGPT, Claude and Google's AI surfaces and found less than a 1 in 100 chance of getting the same list of brands twice, and less than 1 in 1,000 of getting the same order. Its conclusion was that a rank position in an AI answer is not a stable thing to report; a visibility percentage over many prompts, repeated, is.

So ask: does the tool report a single run as truth, or does it tell you a run is a snapshot and show movement over time? The second is more honest, and it is the one that survives a client challenging a number. Are AI visibility scores accurate? goes through what to say when they do.

4. Does it show the sources?

A score says the brand appears in 12% of answers. It does not say why. The sources do: which pages the engine read, and who controls them. If the answers are built from review sites and a competitor's blog, publishing more on your client's own site will not move the number.

A tool that stops at the score leaves you to do this by hand. How to see which sources an AI answer is built from shows the manual version, and it is slow.

5. What can you hand over?

Think about the last day of the month. What does the client receive? A login they will not open, or a document?

Check the format (link, PDF, both), whether it reads without you explaining it, and whether you can put your own name on it and whether that costs extra. Ask to see a real one, not a mock-up.

6. Does it write any of the work?

Some tools stop at diagnosis. Others draft the pages for the client's site or the pitches to third-party sites. Diagnosis-only is a fine product, and cheaper, if you already have a team that writes. If you are a freelancer without spare hours, the drafting is where your time goes. GeoShark drafts both, and you review before anything goes anywhere.

A two-week test

Do not decide from a demo. Pick one real client and write 20 buyer questions for them, spread from "does this kind of service exist" to "best option for my situation". Run them in each tool, and run them again a week later.

Then check three things. Do the two runs tell a coherent story, even if the exact answers moved? Can you see why the brand is missing? Would you be comfortable putting the output in front of the client? The tool that passes all three is the one to keep.

If you want a shorter route, how to track your brand in AI search answers covers the measures you should be comparing on, and GEO vs SEO: what changes for a small agency covers why the sources matter more than the score. New to the term? Start with what GeoShark and GEO are.

FAQ

What should an agency look for in an AI visibility tool?
Six things: how it charges as your client count grows, which engines it covers, how it treats the variation between runs, whether it lists the sources each answer cites, what report you can hand a client, and whether it drafts any of the fixes. The first and the fifth decide most agency purchases.
Should an AI visibility tool charge per brand or per prompt?
It depends on how many questions you track per client. Per-brand pricing is simple but can penalise a roster of small clients or reward a heavy one. Per-prompt pricing follows the work more closely. Multiply the prompts you would track per client by your client count and compare the totals, not the headline prices.
How many AI engines does a tool need to cover?
ChatGPT and Google's AI Overviews carry the most volume, so start there. Perplexity is worth having because it shows its citations, and Gemini is worth having if a client's buyers use Google's assistant. More engines matter less than whether each one is read consistently.
Can I test an AI visibility tool before paying?
Most offer a trial or a sample, and you should use it on a real client. Ask 20 of their buyer questions, run them twice a week apart, and check whether the two runs tell a coherent story. GeoShark's trial runs 14 days with no card.

Sources

  1. AIs are highly inconsistent when recommending brands or products — SparkToro (Rand Fishkin), 2026
  2. GEO: Generative Engine Optimization (Aggarwal et al.) — arXiv, 2024
  • AI visibility tools
  • agencies
  • comparison

New to this? Start with what GeoShark and GEO are.