TL;DR
- 16 KPIs, four layers. Answer: visibility, list position, mention rate, sentiment, share of voice. Sources: self citations, competing cited domains, citation ranking, tracked pages, Reddit and YouTube mentions. Site: crawler hits with response codes, AI assistant visits, site health, content gaps, published pieces. Demand: prompt volume and difficulty.
- Visibility is position plus context minus competitor weight, scored per response, with absences counted as zero.
- Share of voice has two formulas. Mentions divided by all answers is coverage. Mentions divided by all tracked brand mentions is competitive share. Promptwatch reports both.
- A mention names your brand. A citation links to your page. Promptwatch tracks them as separate KPIs.
- Sentiment is one score from 1 to 100 per brand per answer, where 50 is balanced and thin mentions are excluded.
The four layers of AI search measurement
Sixteen KPIs, four layers.
| Layer | KPIs |
|---|---|
| The answer | Brand visibility, list position, mention rate, sentiment, share of voice |
| The sources | Self citations, competing cited domains, citation ranking, tracked pages, offsite mentions on Reddit and YouTube |
| The site | AI crawler hits with response codes, AI assistant visits, site health issues, content gaps, published pieces |
| The demand | Prompt search volume, prompt difficulty |
Every answer layer metric in Promptwatch is sliceable by model, prompt, topic, tag and location. That dimensionality is where the insight lives. A blended visibility score hides that you are strong in ChatGPT and absent from AI Mode, or comfortable in English and invisible in German.
Brand visibility: how much of the answer do you own?
Visibility is a weighted score, not a mention count. Here is how Promptwatch calculates it, per response:
Visibility score = Position + Context − Competitor weight
- Position sets the baseline. A brand named as the lead recommendation or the first list entry scores highest. Lower placements score less, roughly 65 to 80 for second, 50 to 65 for third, and 20 to 45 for brands sitting at the bottom of an answer or mentioned only in passing.
- Context adjusts for how you are described. Explicit recommendation and positive language push the score up. Hedged phrasing and criticism pull it down.
- Competitor weight reduces the score when rivals dominate the same answer.
The overall figure is the mean across responses, and responses where the brand is absent count as zero. That last detail matters more than it looks. Averaging only the answers you appeared in produces a flattering number that rises when your coverage falls, because the weak mentions drop out of the denominator along with the misses.
![visibility over time, brand against two competitors]](https://cdn.sanity.io/images/k0cx5eld/production/8c6f7fb08569bd57c859457bd27d9eb8a45a37a6-2644x986.png)
Two tools will return different visibility scores for the same brand on the same day, and neither is lying. A composite score is a weighting decision. The useful question to put to any vendor is not what your score is but what goes into it.
Mention rate: how often do you show up at all?
Mention rate is the count of answers naming your brand. In Promptwatch it is measured independently rather than derived from the visibility score, which means the two can move apart. Coverage can widen while prominence falls, and that gap is usually the most interesting thing on the chart.
What it does not tell you: anything about position, and anything about whether the mention helped you. Its job is coverage.
Read it at topic or tag level rather than per prompt. A single prompt is a small sample and a noisy one, which is the subject of the section further down. Grouping prompts into monitors by intent is how you get a number that holds still long enough to act on, and it is the first thing worth setting up in prompt tracking on a new account.

List position: where in the answer do you land?
When an answer names several brands, the order is not decorative. Being first or second in a recommendation list is materially different exposure from being eighth, and buyers rarely read past the top few.
Position is more volatile than presence, so it needs more aggregation before it is reportable. It is also the metric most often quoted from a screenshot, which is the worst possible use of it.
Report the distribution rather than the average. An average position of 4 could mean you are consistently fourth, or that you are first half the time and absent the rest. Those are different businesses.
Share of voice: two views, two denominators
Share of voice is used for two different calculations across the category, and both are legitimate. Promptwatch reports both.
| View | Formula | Answers | Moves when |
|---|---|---|---|
| Share of voice trend | Answers mentioning you ÷ all answers in the period × 100, organic prompts only | How much of the conversation you show up in, over time | Your coverage changes |
| Brand comparison | Your mentions ÷ mentions of all tracked brands × 100 | How much of the conversation you own against rivals | Your coverage or a competitor's changes |
The trend view excludes branded prompts. This is deliberate. If you track prompts with your own name in them, your coverage number rises without your visibility improving, and you have built a vanity metric. Filtering to organic prompts keeps it honest. Most tooling does not disclose whether it applies this filter, which is one concrete reason two vendors can report a different share of voice for the same brand on the same day.

Sentiment: what is AI actually saying about you?
Promptwatch scores sentiment as a single reading from 1 to 100 per brand per answer, where 50 is balanced. It is not three polarity buckets, which matters because a 62 and a 38 are genuinely different readings rather than both landing in "neutral". Mentions too thin to judge are excluded, so sentiment is computed over a smaller base than visibility. The two do not share a denominator and should not be read as though they do.
The more useful distinction is between sentiment and accuracy, and the second one is the bigger risk. An answer can be enthusiastically positive and wrong about your pricing. That is a strong sentiment score and a reputational liability on the same response.
Sentiment is also the fastest moving KPI to fix. Training data changes slowly, but the sources shaping how a model talks about you can often be corrected directly, sometimes with a single email to a review site carrying manipulated listings. If you are choosing where to start, tracking how AI platforms describe your brand usually returns faster than chasing coverage.

Citations: which sources are winning, and are any of them yours?
A mention names you. A citation attributes a claim to your page, and it is the only one of the two that can send a click. Report them separately.
Five metrics live here.
- Self citation rate. How often your own domain is the source behind an answer.
- Competing cited domains. Which domains get cited instead of yours on prompts you care about. Operationally this is the most useful number in the layer, because it is a target list rather than a score.
- Citation ranking. Where your citation sits within the answer's source list.
- Tracked pages. Which specific URLs earn citations, and for which prompts. This is what tells you whether the page you built for a topic is the page being read.
- Offsite mentions on Reddit and YouTube. The sources you do not own but can influence, and a large share of what gets cited in practice. Reddit visibility belongs in the source mix report, not only in the content plan.

The site layer: the KPIs that move first
- AI crawler hits per page, with response codes. Hit counts are interesting. Response codes are diagnostic. If GPTBot or ClaudeBot is collecting 403s and 404s on your comparison pages, nothing in the answer layer is interpretable, and you will spend a quarter rewriting content that was never read. Promptwatch picks this up through a CDN or edge integration, which is the same plumbing behind seeing which AI bots crawl your site.
- AI assistant visits. Actual humans arriving from an AI answer, which is a different thing from crawler traffic and routinely confused with it. Agent Analytics separates the two so a spike in bot requests never gets reported as demand.
- Site health issues. Page level problems that block eligibility before anything else can happen.
- Content gaps. Prompts where you are absent or thin. This is the queue everything else feeds into, and it is where Content Agents turns a measurement into a piece of work.
- Published pieces. What you shipped against those gaps, so your trend line has annotations. A visibility chart without publication markers is a chart you cannot explain to anyone.
The funnel underneath all of this is simple: crawled, then cited, then mentioned, then recommended, then clicked. Most teams instrument the last two.

Prompt volume and difficulty: is the prompt worth winning?
Prompt volume estimates how often a question is actually asked. Promptwatch derives it from the keywords attached to a prompt rather than the prompt sentence, taking a weighted average across monthly volumes with AI search weighted five times, Bing one and a half, and Google half. It refreshes at most once a quarter and displays as a band rather than a false precision figure.
Prompt difficulty estimates the effort to become visible, based on the past thirty days. It blends keyword competition in the topic area with the authority of the sources AI models already cite for that prompt.
The practical use is comparative. A 40% mention rate on low difficulty prompts nobody asks is worse than 15% on the prompts that drive your pipeline, and without these two numbers you cannot tell those situations apart. Prompt volumes is the screen to sort by before signing off a content plan.

What to report, and how often
Four KPIs carry most of the reporting weight, because between them they answer four different questions.
| KPI | What it tells you |
|---|---|
| Share of voice | Reach. How much of the conversation you appear in |
| Visibility | Prominence. How much of the answer you own when you do appear |
| Sentiment | Tone. Whether appearing is helping or hurting you |
| Self citations | Sourcing. Whether your own pages are what the model is actually reading |
Weekly is the smallest useful cadence. Prompts run once a day per model in Promptwatch, so a single day is a thin sample that moves on noise, and a daily report will have you chasing changes that were never there.
The unit matters here. A weekly read is reliable at topic or monitor level, where a sixty prompt monitor across three models has already produced more than a thousand answers. An individual prompt still needs the fourteen day window from the section above before it carries a conclusion.
Use weekly internally to catch drops early, while there is still time to trace one back to a crawler error, a competitor's new comparison page, or a source that quietly stopped being cited. Use monthly for the client facing or executive summary, and add crawler activity and content published so the trend line has causes attached to it rather than just movement. Both reports can go out automatically on a schedule, which matters more than it sounds. A report that depends on someone remembering to build it is a report that stops after two months.
Whatever the cadence, every report should state the same six things: the prompt set and whether it changed, the engines, the markets and languages, the aggregation window, the competitor set, and the formula behind any composite score. Without them, this month's number cannot be compared to last month's, let alone to a competitor's.
That disclosure is cheap and almost nobody does it. It is also the difference between a dashboard and evidence, which matters the moment someone asks you to justify the investment. If the people reading your reports live in other tools, Promptwatch data can be pulled into Data Studio alongside GA4 and Search Console so the AI numbers sit next to the ones they already trust.
Frequently asked questions
What are the most important AI search visibility KPIs?
Visibility score, mention rate, list position, sentiment, share of voice, citation rate and self citation rate cover the reporting core. Add AI crawler hits with response codes if you want a leading indicator, because eligibility failures invalidate everything above them.
What is the difference between an AI mention and an AI citation?
A mention names your brand in the answer text. A citation attributes a claim to one of your pages and links to it. You can be mentioned constantly while every citation points at a competitor's blog, which is invisible if you track mentions alone.
How is AI share of voice calculated?
Two ways, both valid. Against all answers in a period it measures your coverage. Against the mentions of all tracked brands it measures your position against rivals. Promptwatch reports the first as a trend and the second on the brand comparison chart.
Why do two AI visibility tools give my brand different scores?
Because a composite score is a weighting decision. Different prompt sets, position weights, run counts, branded prompt filters and source rules produce different numbers for the same brand on the same day. Ask for the formula before comparing scores.
