Two years ago nobody sold this. Today there is a category, the tools cost real money, and the search term ai visibility tools went from a rounding error to roughly 1,600 US searches a month, with advertisers bidding around fifty dollars a click.
That combination, new category and expensive clicks, produces a lot of vague marketing. This piece is about what these tools actually measure, which of those measurements are worth paying for, and what the numbers cannot tell you.
Disclosure: I am building an AI visibility module inside my own tool, and it is not finished. Nothing below is a pitch for it. Where my own product is relevant I have said so and said plainly what it does not yet do.
Why the category appeared at all
A search results page has room for everyone. Ten organic slots, a map pack, some ads, and a second page underneath. Being eighth is worth less than being first, but it is not worth nothing.
A generated answer has room for two or three. When somebody asks an assistant to recommend a plumber in their city, or the best project management tool for a small team, the answer names a handful and stops. There is no page two. If you are not named, the person never learns you exist, and nothing in your analytics records the loss.
Classic rank tracking cannot see this. Your positions can be perfect while a growing share of the demand never reaches a results page at all.
The five things worth measuring
Strip away the dashboards and every tool in this category is doing the same basic job: running a set of prompts through one or more assistants, reading the answers, and extracting structured facts from unstructured text. The useful facts are these.
1. Mention
Were you named, yes or no. This is the floor. It is also the only metric that is genuinely unambiguous, which makes it the one to anchor reporting on.
2. Position within the answer
If three companies are named, being first is not the same as being third. Assistants tend to elaborate on the first name and list the rest, so ordering carries real weight even though there is no numbered list to point at.
3. Sentiment and framing
Being mentioned is not automatically good. A cheaper option if you can live with a thinner feature set is a mention. So is the one most agencies settle on. A tool that reports mention count without framing is reporting half the story.
This is also the softest number in the category. Sentiment classification on two clauses of text is genuinely hard, and I would treat any vendor’s sentiment score as directional rather than precise.
4. Citations
Most assistants show their sources. Those source lists are the most actionable output in the whole category, because they tell you which pages the model trusts on your topic. If the same three industry blogs get cited for every prompt in your category, you now have an outreach list that was built for you.
Worth noting: the cited page is frequently not the brand’s own site. It is a comparison article, a forum thread, or a directory. That is a content strategy insight, not a ranking problem.
5. Query fan-out
Before an assistant writes, it usually runs its own searches. Some platforms expose those queries. They are often not the phrasing anyone would type, and almost nobody is optimising for them yet. Of the five metrics here this is the one with the widest gap between how useful it is and how much attention it gets.
Share of voice, and why it needs a denominator
Every tool in this space sells a share of voice number. The number is only meaningful if you know what it is a share of.
| What it can mean | What it tells you |
|---|---|
| Percent of your prompts where you appear | Coverage of a topic you chose |
| Your mentions as a share of all brand mentions | Position relative to named competitors |
| Weighted by position in the answer | Closer to attention, harder to audit |
All three are legitimate. They are not comparable to each other, and a number quoted without its definition is not a number. Ask which one a tool uses before you put it in a client report.
There is also a self-inflicted problem worth naming: you write the prompts. Write flattering prompts and your share of voice looks excellent. The metric only means something if the prompt set reflects how customers actually ask, which usually means it should include prompts you expect to lose.
Why one API for every model is a shortcut worth avoiding
Some tools route every request through a single aggregator. It is cheaper to build and it produces a tidier dashboard.
The problem is that different assistants use different retrieval systems. What ChatGPT surfaces, what Gemini surfaces, and what appears in a Google AI Overview are the products of separate pipelines with separate indexes and separate freshness. A number blended across them describes no system that exists.
If a vendor cannot tell you which surfaces they query directly, the comparison across models is decorative.
What these numbers cannot tell you
They are not deterministic. Ask the same model the same question twice and you can get two different answers. Any single reading is noise. Only the trend across repeated runs means anything, which is why cadence matters more than precision here.
They are not traffic. A mention is not a visit, and most assistants send very little referral traffic even when they name you. The value is in the recommendation, not the click, which makes attribution genuinely unresolved. Nobody in this category has solved that yet, and vendors who imply they have are overselling.
They do not travel across languages reliably. Coverage for prompts outside English varies by provider and by model, and it is currently the least documented part of the whole category. If you work in a non-English market, make this the first question you ask, not the last.
What to ask before you buy one
- Which assistants do you query directly, and which are inferred?
- How is share of voice calculated, precisely?
- How many times is each prompt run before a result is reported?
- Do you return the citation list, or only the mention?
- What is the per-run cost at my prompt count, and is it capped?
- What is your coverage for the languages I actually sell in?
Six questions. A vendor who answers all six in plain numbers is doing the work. A vendor who answers in adjectives is selling a dashboard.
Where this sits for local businesses
Most of the tools in this category are built for national brands tracking category-level prompts. For a business that serves one city, the interesting prompts are local, and local prompts inherit the location problem that already plagues rank tracking: the answer depends on where the assistant thinks the person is.
That makes local AI visibility a harder measurement problem than the national version, not an easier one. It is the part of this category I find least solved, and it is why the AI visibility module starts from the same coordinate discipline as the rest of the tool rather than treating AI as a separate product.
Common questions
How often should prompts be run?
Weekly is enough for most categories, because answers move slowly and each run costs money. Daily runs mostly buy you noise. What matters is that the cadence is consistent, so the trend line means something.
How many prompts do I need?
Twenty to thirty covering a real topic is more useful than two hundred covering everything. The set should include prompts you expect to lose, or the number is flattery.
Can I influence what an assistant says about me?
Indirectly. The lever is the source material the model retrieves, which means the pages that get cited in your category. That is closer to digital PR and comparison content than to traditional on-page work.
Is this replacing rank tracking?
Not yet, and probably not entirely. Classic results still carry most local commercial intent. Treat AI visibility as a second surface to measure, not a substitute for the first.
What does it cost to run?
The underlying data is bought per request. As a rough anchor, twenty-five prompts across three models sits around one euro per run at current provider rates, before whatever a tool charges on top. Ask any vendor to show you that split.