The denominator changes the result: the hidden trap in AI KPIs
The same observations can produce very different percentages. Learn the right formulas and denominators for AI visibility, recommendation rate, coverage and share of voice.
A KPI is not only defined by the counter. The denominator determines the question the percentage is answering. In measuring AI visibility, the same nine recommendations can become 32.1%, 9.8%, or some other value without any calculation being arithmetically wrong.
The issue occurs when the dashboard displays the percentage but hides the cohort, unit of analysis, failed responses, or inclusion rule. Then we compare numbers that look the same and measure different things.
Same dataset, five KPIs
We use one consistent example:
- 100 planned interactions;
- 92 eligible and analyzable responses;
- 28 responses where the brand appears;
- 9 answers where the brand is recommended;
- 41 target brand mentions;
- 123 mentions of all brands;
- 12 occurrences of first-party citations;
- 80 occurrences of all citations.
| KPI | Formula | Result | Question answered |
|---|---|---|---|
| AI visibility | 28 ÷ 92 | 30.4% | How many eligible responses does the brand appear in? |
| Recommendation rate among appearances | 9 ÷ 28 | 32.1% | When the brand appears, how often is it recommended? |
| Recommendation coverage | 9 ÷ 92 | 9.8% | In how many eligible responses is it recommended? |
| Mention share of voice | 41 ÷ 123 | 33.3% | What part of all mentions belongs to the brand? |
| First-party citation occurrence share | 12 ÷ 80 | 15.0% | What fraction of citation occurrences lead to own sources? |
The numbers do not contradict each other. Each describes a different stage. The error would be to present the 32.1% recommendation rate as if the brand was recommended in 32.1% of all responses.
Denominator for AI visibility
The convention used in this cluster is:
AI visibility = eligible responses in which the brand appears ÷ all eligible responses that can be analyzed.
“Eligible” must be defined before analysis. A scenario about a category in which the company does not operate should not be included merely to change the score. At the same time, we cannot select only the scenarios where we expect the brand to perform well.
Tests where the brand name is entered explicitly may be useful for measuring perception and factual accuracy, but they do not demonstrate organic discovery. Therefore, branded tests should not inflate the organic visibility KPI.
I explain the full protocol in What AI visibility is and how to measure it correctly.
Recommendation rate has at least two legitimate denominators
The question "how often is the brand recommended?" can mean:
- among the responses where the brand appears — 9/28 = 32.1%;
- of all eligible responses — 9/92 = 9.8%.
The first value shows the brand's ability to turn an appearance into a recommendation. The second shows the coverage of the recommendation in the entire tested space.
Both are useful. None should be labeled simply "AI recommendation 32%" without full name and formula.
Share of voice: conversations, mentions or recommendations?
Share of voice can use different units:
- number of conversations in which each brand appears;
- total number of mentions;
- number of recommendations;
- a measure of prominence or position.
If a response mentions the same brand five times, occurrence SOV assigns it five events. Conversation SOV assigns it one. The choice depends on the business question, but the report must declare the unit.
Competitive set matters just as much. If we add or remove competitors, the share changes even if the number of brand mentions remains the same.
Citations: URLs, domains, occurrences or conversations
"15% first-party citation rate" can hide at least four formulas:
- 12 first-party occurrences out of 80 citation occurrences;
- number of unique first-party URLs out of all unique URLs;
- conversations with at least one first-party quote from all quoted conversations;
- first-party quote conversations from all eligible conversations.
We cannot mix occurrence-level with conversation-level. A single response with five links to the same domain looks very different in the two methods.
The article on query fan-out explains why a page can be found in an intermediate search without being cited in the final response.
Factual accuracy must not turn the unknown into error
For brand claims we need at least three states:
- correct — supported by source of truth and current;
- incorrect — contradicted by evidence;
- not confirmed/not grounded — we don't have enough evidence.
If we call all unconfirmed claims “false”, the accuracy rate becomes artificially harsh. If we remove them from the denominator without reporting them, we hide a knowledge-base coverage problem.
The eight failed responses do not disappear
In the example, 100 interactions were planned, but only 92 produced analyzable responses. Visibility is calculated on those 92, while the other 8 are reported separately as failures.
It is not acceptable to remove failures and only show 30.4% without "92/100 parseable responses". If failures are concentrated in a particular engine, language, or script type, they can skew the comparison.
A useful rule is:
Any percentage must be accompanied by n and coverage. Any segment too small should be marked "insufficient data".
Micro-average vs macro-average
Assume we measure two engines:
- Engine A: 20 occurrences out of 80 responses — 25%;
- Engine B: 8 occurrences out of 12 responses — 66.7%.
The simple average of the percentages is 45.8%. Aggregation of all observations is 28/92 = 30.4%. The first method gives equal weight to each engine; the second gives weight to each observation.
None is automatically correct. We have to choose between:
- macro-average — each engine, country or segment has equal weight;
- micro-average — each observation has equal weight;
- weighted average — weights reflect a justified distribution.
The dashboard must show segmented values before the aggregate total.
A denominator change must be treated as a series break
A historical series becomes misleading if the formula changes silently. Let's say that in July we include only completed responses, and in August we include interrupted responses in the denominator. The score may drop even if the number of brand appearances remains the same.
The same problem occurs when we expand the list of scenarios, add a language, change the interface of an engine, or reclassify what "recommendation" means. These are not simple dashboard updates. There are changes to the measuring tool.
For every change we keep two things:
- the exact date from which the new definition takes effect;
- an overlap period where we calculate the old and new formula on the same observations.
If history recalculation is possible, we label the methodology version. If not possible, we mark the break and do not draw a solid line over it.
Example of an auditable record
Instead of "visibility 30.4%", a full operational sheet might say: "28 of 92 eligible responses mentioned the brand; 100 interactions were planned; 8 did not produce an analyzable result; cohort contains non-branded scenarios in Romanian; period is July 15–31; Wilson 95% interval is 22.0%–40.5%”.
The wording is longer, but it allows an analyst to reproduce the number and a decision maker to understand its limits. The dashboard can keep the percentage as the main visual element, provided that all this data is accessible in a tooltip, export or methodological registry.
At the decision level, the action threshold must be set before the outcome. For example, we can investigate a segment if the decline persists in two comparable waves and exceeds the estimated uncertainty. Choosing the rule after seeing the graph favors convenient conclusions.
Confidence interval: percentage has margins
28/92 produces an estimate of 30.4%. A Wilson 95% interval is approximately 22.0%–40.5%. For 9/28, the 32.1% estimate has a much wider range: about 17.9%–50.7%.
NIST describes the Wilson method for ratio ranges. It is more informative than displaying a percentage with two decimal places and no uncertainty.
There is also a limit: repeated observations on the same scenarios may not be completely independent. A simple binomial interval may seem more precise than it is. That's why we also keep the structure of the cohorts, not just the total.
First-party data and external measurement are not the same thing
In 2026, Google announced Search Console reports dedicated to visibility in generative features, with impressions, pages, countries, devices and dates. In the original documentation, the rollout was limited to a subset of sites.
A first-party impression reported by Google is different from a simulated conversation in an external tool. Both can be useful, but should not be lumped into a single unlabeled KPI.
Google says Search Console is the source of truth for Google Search performance and Analytics for site behavior. The same discipline must be maintained here: the source determines what we can assert.
One minimum contract for any AI KPI
Next to each indicator there should be:
- the full name of the metric;
- formula;
- numerator and denominator;
- unit: conversation, turn, mention, URL or citation;
- n and coverage;
- engine, interface, language and geography;
- cohort and period;
- confidence interval;
- exclusions;
- series breaks.
What I will not do in AYSA reports
The proposed editorial methodology for AYSA refuses some shortcuts:
- we will not publish universal "good/poor" thresholds without a reference cohort;
- we will not hide failed answers;
- we will not call a mention a recommendation;
- we will not compare engines with incompatible samples without warning;
- we will not present a before/after difference as demonstrated causation;
- we will not assert the existence of a feature before product verification.
These are guidelines for methodology and transparency, not a statement that all reports and integrations are already available in the product.
Conclusion
A percentage is not the truth about a brand. It is the result of a formula applied to a set of observations.
Before asking whether 32% is good, we need to know whether it means recommendations among appearances, recommendations among all conversations, share of voice or something else. Then we check n, coverage, cohort and uncertainty.
The denominator is not a technical note. It is the definition of the business question.
Sources
- NIST/SEMATECH — Confidence intervals for proportions
- Google Search Central — Generative AI performance reports
- Google Search Central — Using Search Console and Analytics data
Sources last checked: 31 July 2026. Calculations in the example are reproduced from the numerators and denominators shown.