TL;DR
- An AI visibility score is a 0–100 measure of how prominently your brand appears in AI-generated answers, not how often it's named.
- The Promptwatch formula, in full: per response,
Position + Context − Competitors. Overall, the plain average of every response, with absences counted as 0. - Mentions and visibility are different things. A model saying "I'm not familiar with Acme" has technically mentioned you.
- Your score will differ in every tool you try. Different prompt sets, engines, run counts and absence handling. Compare within one tool over time, never across two.
- One check isn't a score. Identical prompts re-run minutes apart returned overlapping sources only 32–43% of the time in University of St. Gallen testing.
What Is an AI Visibility Score?
An AI visibility score is a 0–100 metric measuring how prominently a brand appears in AI-generated answers across engines like ChatGPT, Gemini, Perplexity and Google AI Overviews. It combines whether the brand appears, where in the answer it appears, how it's described, and how much of the answer competitors take.
The shift it captures is simple. In SEO, the unit of measurement is the page: you rank at a position, and the position is a number. In AI search, the unit is the response. You're either inside a generated paragraph or you're not, whether as the lead recommendation, the fourth name in a list, or a brand the model admits it hasn't heard of.
A score exists because nobody can run a channel on a yes/no answer repeated four thousand times.
How an AI visibility score is calculated
Every credible score uses four inputs:
- Presence. Does the brand appear in the response at all?
- Position. Where in the answer? Lead recommendation, mid-list, or a passing aside.
- Framing. Is the mention a recommendation, a neutral fact, or a caveat?
- Competition. How many rivals share the answer, and how much of it do they take?
Those get scored per response, then averaged into one number. The differences between tools come from how each step is handled, and from one step in particular.

The input that decides everything: the answers you're not in
The biggest driver of divergence between two scores for the same brand isn't weighting. It's whether responses that never mention the brand are counted at all.
| Method A: average of mentions only | Method B: absences counted as 0 | |
|---|---|---|
| Responses analysed | 10 | 10 |
| Responses mentioning the brand | 4 | 4 |
| Scores on those four | 95, 72, 88, 65 | 95, 72, 88, 65 |
| Remaining six responses | Excluded | Counted as 0 |
| Reported score | 80 | 32 |
Same brand, same week, same underlying data. Method A measures how well you do when you show up, which flatters relentlessly. The answers where you're invisible are exactly the ones that should worry you. We use Method B.
How our Visibility Score is calculated
Per response: Position + Context − Competitors. Overall: the average of every response scored.
Position sets the baseline
| Where the brand sits in the answer | Baseline |
|---|---|
| Lead recommendation or first list entry | Highest |
| Second | 65–80 |
| Third | 50–65 |
| Buried at the bottom, or mentioned in passing | 20–45 |
| Not mentioned | 0 |
Context adjusts it
Context is how the answer talks about you. Strong recommendations and positive framing ("a great choice", "highly recommended", "best in class") push the score up. Negative framing ("avoid", "limited features", "not reliable") or hedging ("might work", "I don't have much information on…") pull it down.
Competitor weight pulls it down
Being named second in a two-brand answer is a stronger position than being named second in a twelve-brand list. The score reflects that.
Here's a real answer to "What's the best project management tool?":
When looking for the best project management tools, Acme stands out as the industry leader with comprehensive features. Other options include Competitor A, which offers basic functionality, and Competitor B for smaller teams.
Acme leads the answer, is framed as the category leader, and the two competitors are handled as secondary options. High position, positive context, light competitor weight. This response scores in the nineties.
Then average across every response
Σ all response scores ÷ total responses analysed
So 95 + 72 + 88 + 0 + 65 + … = 85. The zero in that sequence is a response where the brand didn't appear, and it counts.
Why not just count mentions?
Because a raw count is misleading in one specific, common way: models frequently name a brand only to disclaim it. "I'm not familiar with Acme" contains the brand name and registers as a mention in any string-matching system.
Under the Visibility Score, those responses get a near-zero baseline. Which means a brand can be technically mentioned in 100% of responses and still score very low. That gap is the whole argument for using a score rather than a tally.
Sentiment and citations sit outside the formula deliberately. Sentiment is scored separately, 1–100, based on how the answer talks about you rather than how prominently you feature. An answer devoted entirely to problems with your product is high visibility and low sentiment at the same time, and folding the two together destroys exactly the signal you needed. Citations are tracked separately too, because being named and being the source are different achievements.
Why your score is different in every tool
Run the same brand through four platforms and you'll get four numbers. Five reasons why:
- Different prompt sets. The largest factor, and the one you control. Peec AI's analysis of 37,804 responses across five engines found that prompt format alone moves visibility: ranking-style prompts surfaced roughly 20% more brand mentions than open-ended ones, and concise keyword-style prompts up to 25% more than conversational framings. Change the shape of your prompt list and your score moves without your brand changing at all.
- Different engines. Source overlap between engines is low. A four-engine score and a seven-engine score aren't measuring the same market.
- Different run counts. More on this below.
- Different absence handling. The 80-versus-32 problem above.
- Different normalisation. Some scores are absolute, some are indexed against a competitor set. An indexed score can fall while your real presence grows, purely because a rival improved faster.
The practical rule: compare within one tool over time, and never compare a number from one tool against a number from another.
Most AI visibility platforms don't publish their methodology, so there's no way to reconcile the two anyway.
What's a good AI visibility score?
There's no universal benchmark, and anyone offering one is describing their own product. The bands below circulate widely and are useful as rough orientation, not as a standard.
| Score | Usually means |
|---|---|
| 0–8 | Not being recommended; entity signals weak or absent |
| 8–25 | Appears inconsistently, rarely as a primary recommendation |
| 25–50 | Regularly on AI shortlists for competitive queries |
| 50–75 | A default answer for many category questions |
| 75–100 | The primary recommendation; rivals mentioned secondarily |
Competitive density dominates: reaching 50 in CRM is a different achievement to reaching 50 in a niche vertical with three rivals. And the score is a leading indicator. There's typically a lag of a couple of months between visibility improving and attributable pipeline showing up, which is how teams talk themselves out of the metric too early.
The only benchmark that means anything is your own competitor set, on your own prompt set, tracked over time.
How to improve your AI visibility score
- Fix absence before prominence. Moving a response from 0 to 45 does more for the average than moving one from 65 to 80. Start with the prompts where you don't appear at all.
- Earn the position, not just the mention. Comparison-shaped, list-shaped, explicitly recommendable content is what gets pulled to the top of a generated answer.
- Fix how you're described. Hedged and negative framing hits the context component directly. Inconsistent brand descriptions across your site, LinkedIn, G2 and directory listings are the usual cause.
- Check that the crawlers can get in. Rate limits, 403s and blocked user agents cap everything above them. Agent Analytics shows which AI bots are actually reaching your pages.
- Work the third-party sources. Much of what feeds these answers isn't your site at all. It's review platforms, forums and community threads.
Then close the loop: find the prompts where you're absent and turn them into pages.
See your own score
The free AI Brand Visibility Report gives you a baseline across engines with no account needed. Continuous tracking with your own prompt set, competitor benchmarking and per-market breakdowns is the paid product.
Frequently asked questions
Is an AI visibility score the same as share of voice?
No. Share of voice is your proportion of the conversation relative to named competitors. Visibility is your absolute prominence across all analysed responses. You can hold a large share of a conversation that barely happens.
Can a brand be mentioned in 100% of responses and still score low?
Yes. Responses that name your brand only to say the model isn't familiar with it score near zero. This is the clearest reason to use a score rather than a mention count.
Why did my score drop overnight when I changed nothing?
Models re-crawl and re-generate constantly, and a single response can shift on its own. Competitor movement and model updates both register. Track the trend over weeks.
Does my AI visibility score affect my Google rankings?
Not directly. But entity clarity, structured content, third-party validation and citable original data feed both, and rising AI visibility tends to lift branded search volume.
Can I combine several brand names into one score?
Yes. Brand merging rolls multiple entities or IP into a single brand's visibility tracking and reporting.
