SEO, GEO & AI

What AI visibility is and how to measure it correctly

What it means to be visible in ChatGPT, Gemini, Claude or Perplexity, how to choose the denominator, and why one prompt cannot produce a credible score.

Dashboard for measuring a brand's visibility in AI answers, including sample size, coverage and confidence interval
An AI visibility score becomes useful only when we know which conversations, engines, languages and measurement rules sit behind it.

AI Visibility is the proportion of eligible interactions where a brand, product or other entity appears in a response generated by an AI engine. The percentage cannot be properly interpreted without the question set, engine and interface tested, language, geography, period, number of repetitions, detection rules and denominator used.

In other words, “the brand has 30% AI visibility” is not a complete conclusion. It is the start of a question: 30% of what?

In this guide I separate visibility from mention, citation, and recommendation, explain the formula I use in this series, and show why a single prompt cannot support a credible report.

What AI visibility means

In traditional search, visibility is usually associated with rankings, impressions and appearances in a results list. An AI answer, however, is not a stable list of ten links. It may synthesize several sources, compare options, mention a brand without recommending it, or make a recommendation without citing that brand’s website.

Therefore, I treat visibility as an observed event in a cohort of interactions:

The brand appears or does not appear in an eligible response, according to rules declared before analysis.

"Appears" must be defined. It can mean the official name, a product name, an abbreviation or a known variant. A serious system keeps a registry of aliases and avoids ambiguous matches. "Orange", for example, can be a company, a color or a fruit. Simple detection of a string is not always enough.

Visibility is not ranking

A conversational engine can produce different answers for the same question. The result may vary depending on the time, model, interface, language, location, conversation history, available sources and the wording of the request.

There is no single, permanent position that we check once a month. We can measure the frequency and form of occurrence in a repeatable protocol, but we shouldn't turn the result into a "3rd place in ChatGPT" without explaining exactly how it was obtained.

Google states that AI Overviews and AI Mode may use different models and techniques, and the answers and links displayed may vary. The company also says that compliance with technical requirements and best practices does not guarantee that a page will be indexed or displayed. Similarly, OpenAI explains that a public site may appear in ChatGPT Search and recommends access for OAI-SearchBot, but technically access is not a promise of appearance.

Mention, recommendation and citation: three different events

A common mistake is to call any occurrence a "recommendation". Let's take three hypothetical answers:

  1. "There is AYSA, Company B and Company C in the market." — mention.
  2. "For the situation described, I would consider AYSA because..." — recommendation.
  3. "According to the methodology published by AYSA..." followed by a link — citation.

The same response can contain all three events, only one, or none. Visibility answers the question "does the brand appear?", not the questions "is it chosen?", "is it used as a source?" or "is it favorably described?".

For a business decision, these dimensions should be reported separately:

  • visibility — how often the entity appears;
  • recommendation — how often it is proposed as a choice;
  • citation — how often it is used or shown as a source;
  • perception — what statements and attributes are associated with it;
  • accuracy — how many of the statements are correct and current;
  • impact — what traffic, leads or observable results follow.

The basic formula

For this editorial series I use the following convention:

AI visibility rate = eligible interactions in which the brand appears ÷ total eligible interactions × 100.

It is a stated measurement convention, not a universal standard implemented identically by all tools. The differences appear mostly in the word "eligible".

What is an eligible interaction

An interaction should only enter the denominator if it is part of the established cohort and can reasonably answer the measured question. The protocol must decide in advance:

  • what themes, personas and scenarios are included;
  • which engine and interface are tested;
  • what language and geography are used;
  • how many iterations are executed;
  • how rejections, timeouts and incomplete responses are handled;
  • if a conversation is analyzed in full or on a per-turn basis;
  • which tests are informational and which are commercial or recommendation-oriented.

If we include questions about areas where the brand is not active, the score is artificially lowered. If we only choose situations where we know it occurs, the score artificially increases. The scenario set is part of the measurement, not a hidden detail in the platform.

Why Branded Tests Shouldn't Inflate Organic Visibility

The question “What do you know about AYSA?” is useful for analysing perception and checking facts. It is not, however, an organic measure of brand discovery: the evaluator introduces the name.

Unbranded questions are more suitable for organic visibility, for example: “Which solutions can analyse and apply SEO changes to a WordPress website with human approval?” If AYSA appears, we have a signal of discovery or association with the need. If the question already contains the name AYSA, we may analyse the description, accuracy or perception, but should not count the appearance as organic discovery.

Commercial methodologies differ. Therefore, before we compare two scores we need to check that both exclude branded tests and use the same type of conversations.

A transparent calculated example

Assume we prepare 100 eligible interactions for a brand:

  • 92 produce parseable responses;
  • 8 fail or are excluded and reported separately;
  • the brand appears in 28 of the 92 responses;
  • it is recommended in 9 of the 28 responses in which it appears.
Metric Calculation Result Question Answered
Visibility 28 ÷ 92 30.4% In how many eligible responses does the brand appear?
Recommendation rate among appearances 9 ÷ 28 32.1% When it appears, how often is it recommended?
Recommendation Coverage 9 ÷ 92 9.8% How many eligible responses actually reach the recommendation?

All three percentages are correct, but answer different questions. If a dashboard shows "32.1% recommendation rate" without a denominator, the reader may mistakenly assume that the brand was recommended in nearly a third of all interactions.

A percentage without uncertainty seems more accurate than it is

In our example, the visibility estimate is 30.4%. For 28 occurrences out of 92 observations, a Wilson 95% interval is approximately 22.0%–40.5%. The range is wide because the sample is relatively small and the results are variable.

For recommendation in 9 of 28 occurrences, the estimate is 32.1%, but the approximate range is 17.9%–50.7%. That doesn't make the measurement useless. It just tells us that we are not allowed to present 32.1% as an exact constant.

A credible report shows at least:

  • the estimate;
  • number of observations;
  • confidence interval;
  • response coverage;
  • methodology or model changes.

Why a single prompt is not a measurement

A screenshot can demonstrate that a particular response existed. It does not demonstrate how frequently the brand appears, how stable the result is or how the engine behaves in general.

For a useful measurement we need controlled variation:

  • more relevant scenarios;
  • documented phrasings and personas;
  • repeats under comparable conditions;
  • same detection rule;
  • separation of engines and interfaces;
  • a declared observation period.

Multi-turn conversations add another layer. A brand may appear in the initial response, but disappear after the user specifies their budget, country, required integration, or compliance criteria. First turn visibility and survival to final recommendation are not the same metric.

In addition, some systems may rephrase the question and search for subtopics before composing the answer. I explained this mechanism separately in the guide on query fan-out and hidden queries in AI Search.

Engine, interface, language and geography are part of the result

“I tested GPT” is insufficient. We need to know whether the test used ChatGPT in a browser, an API integration, a mode with web search or another surface. Interfaces may use different tools, sources and behaviours.

Language changes available sources and entity formulation. Geography may change local recommendations, stores, prices, availability or applicable rules. A Romanian brand can be visible in a conversation in Romanian and absent in the same theme formulated in German.

That is why I would not aggregate all engines and countries into one percentage before seeing the separate results. An average can hide the fact that the brand performs well in one market and hardly at all in another.

Coverage, errors and "insufficient data"

If 40% of the tests failed, a score calculated from only the rest of the responses may be technically correct but operationally poor. The report must show coverage: the proportion of planned interactions that produced analyzable results.

Thresholds for insufficient data must also be defined. Not every cell in a dashboard deserves a percentage. For an engine, language or scenario with too few observations, the most honest value may be “insufficient data”.

What technical accessibility can demonstrate

A page blocked by robots.txt, inaccessible to the crawler, or unavailable in HTML has a real problem. But unblocking does not guarantee citation or recommendation.

Google Search Central explicitly says that AI Overviews and AI Mode do not require a special AI file or special schema. The basics matter: indexing, snippet eligibility, internal linking, accessible textual content, good experience, and structured data that matches the visible page.

OpenAI recommends that sites that want to be discovered and cited in ChatGPT Search do not block OAI-SearchBot. And we're talking eligibility and access here, not a guarantee.

For the technical side, I have published the guide separately Your site looks good. But can Google, ChatGPT and AI agents read it?.

How to read an AI visibility dashboard

Before making a decision based on your score, ask these questions:

  1. What is the exact formula?
  2. What is the denominator?
  3. How many observations are there?
  4. What engines and interfaces were tested?
  5. What languages and geographies are included?
  6. Are the scenarios branded or unbranded?
  7. How many runs failed?
  8. How are brand aliases detected?
  9. What is the confidence interval?
  10. Has the model or methodology changed in the period compared?

If this information is missing, the score can be useful as a guiding signal, but it is not yet sufficient to support an important decision.

Where AYSA fits into this methodology

The product thesis on which I am building AYSA is broader than a dashboard: measure → explain → approve → execute → remeasure.

At this stage, the reporting formula and requirements described in the article represent the proposed editorial methodology. They should not be interpreted as a statement that every engine, integration or flow is already available in the product. Commercial pages will explicitly separate available, beta, planned, and methodology-only features.

This separation matters. A measurement product does not become more credible if it presents its roadmap as released functionality.

Verdict

Visibility in AI is not a fixed position nor a permanent property of a brand. It is an estimate obtained from a stated set of interactions, in a given context and over a given period.

A good score doesn't start with a percentage. It starts with a methodology we can inspect: questions, cohort, engine, interface, language, geography, repetitions, denominator, coverage, and uncertainty.

Only then can we ask something really useful: where does the brand appear, where is it missing, when is it recommended, what sources support it, and what intervention is worth testing.

Sources and methodology

Last source check: 31 July 2026. Formulas and intervals in the example use 92 eligible responses and Wilson's 95% interval.