SEO, GEO & AI

What sources AI engines use when they talk about a brand

A page that is found is not automatically cited, and a mention does not reveal its source. A practical taxonomy for first-party, independent and connected sources.

Different web sources passing through selection before appearing in a cited AI answer
There are several steps between a published page and a visible citation. Not all of them can be observed from the outside.

Just because an AI engine mentions a brand doesn't prove what source it used. And just because a page is found doesn't mean it will be cited. Between publication and response there is a sequence of selections: discovery, search, retrieval, processing, supporting a claim, and showing a source.

Some stages are visible. We can see citations, sources displayed and referral traffic. Others are internal and should not be reconstructed from guesswork. A credible methodology separates what we observe from what we merely infer.

Seven different states of a source

  1. Discoverable: the page can be accessed and, where applicable, indexed.
  2. Candidate: The URL appears in a result set for an intermediate search.
  3. Retrieved: the system retrieves the relevant content or fragment.
  4. Used for grounding: the content enters the context supporting the response.
  5. Reflected: an assertion from the source is represented in the response.
  6. Cited: the user sees a link or indicator associated with the source.
  7. Accessed: the user clicks the link and reaches the site.

We cannot assume that all systems go through the same steps or expose the same telemetry. The list is an audit taxonomy, not a universal description of the internal architecture.

What “cited”, “mentioned” and “fetched” mean

Cited describes an observable fact: the answer displays a link to a page or source. Mentioned means that the brand, domain or claim appears in the text. The two events can occur separately.

Fetched or retrieved describes a technical step. In an API there can be a structured result that shows which URLs were returned. In a public interface, a visible citation does not necessarily expose all consulted pages, and a consulted page need not necessarily appear in the response.

The Perplexity documentation explicitly separates the search component from the generation component and recommends using the structured result field for actual URLs, not generating addresses from text. It's a useful example of the difference between system-provided provenance and model-produced text.

Source types for a brand

Type Examples Suitable role Limit
First-party official site, documentation, help center, own data specifications, policies, availability, official position does not independently validate quality
Authoritative regulator, institution, standard, legislation official obligations, definitions and requirements may not evaluate the actual product
Independent press, research, analysis, verifiable review context, comparison and external validation methodology and timeliness differ
Community forums, discussions and user experiences real-world problems, language and borderline cases identity and representativeness may be uncertain
Market participant category pages and trade comparisons positioning and criteria used in the category commercial interest and biased selection
Connected or private files, intranet, authorized applications custom user responses does not represent public organic visibility

The right source depends on the claim. A vendor's website is suitable for its list of declared integrations. An authoritative source is appropriate for a legal obligation. An independent evaluation can support claims about user experience. None automatically replaces the others.

First-party does not automatically mean untrusted

For company-controlled facts — current price, API documentation, return policy, regions served — the official page is often the correct primary source. The problem arises when the same page is used as independent evidence for "best", "safest" or "market leader".

A brand should publish accurate, dated and verifiable information: specifications, definitions, limitations, examples, methods and update dates. Google recommends unique, valuable content for generative search features and explains that they rely on Search's ranking and quality systems. The official guide, however, does not promise inclusion merely because a page was published.

Authority and independence answer other questions

For a regulated claim, an authoritative source takes precedence. For an experience comparison, we need both methodology and independence. For reputation, several independent observations are more useful than one promotional phrase.

A solid response can combine sources. For example, it uses a product's official documentation for capabilities, an institution for legal framework, and an independent analysis for tradeoffs. Diversity is not the mechanical gathering of many domains, but matching the source with the claim.

Romanian example: RO e-Factura

The question “Which solution should I choose for RO e-Factura?” contains at least three subquestions:

  • what obligations and terms apply;
  • what features does each product declare;
  • how the product works in real use.

For obligations, the answer should start from official sources such as the ANAF guidelines and the framework described by the European Commission. For integration, the manufacturer's documentation is the direct source. Independent and dated observations are required for support and usability.

If an engine cites only a commercial page for a tax obligation, that does not automatically make the claim wrong. The source ledger can, however, mark it as a poor source fit and require verification against an authoritative source.

How query fan-out affects selection

A question can be broken down into several searches: definition, topicality, integration, price, risk or geography. OpenAI explains that Search can rewrite the request into one or more targeted queries. A page can become a candidate for one of these queries without being cited in the final response.

This is why the article on query fan-out and hidden queries should be read alongside this guide. Source visibility is not limited to an exact match with the original prompt.

Crawl, indexing and citation are not promises

Crawler access is a technical condition, not a selection guarantee. The OpenAI FAQ for Publishers states that enabling OAI-SearchBot helps content to be discovered, summarized and cited, but the documentation on ChatGPT Search explicitly states that there is no way to guarantee top ranking.

Similarly, indexing in a search engine does not guarantee that a page will be selected for a generative function. Relevance, quality, originality, freshness and other signals can come into play. We do not attribute the power to produce citations to a single file, bookmark, or setting.

The cited source may not be the original source

An assertion can propagate through multiple publications. A secondary article summarizes a study, a commercial page picks up the abstract, and the AI response cites the last page. The link is real, but the origin of the idea is older.

For important claims, we assess the source's citation standing: is the page the primary source, an authorized reproduction, an independent synthesis or merely a repetition? If it reports a percentage, we check whether it links to the study and identifies the period, sample and methodology. If it describes a law, we trace the claim to the institution or current legal text.

This check prevents two errors. The first is to attribute originality to the domain that was accidentally cited. The second is to assume that a phrase becomes true because it appears on many pages that copy each other.

Freshness must be evaluated at claim level

Page date is not enough. An article updated yesterday may retain an old specification, and a page published two years ago may contain a stable definition. The ledger should note the relevant date for the concrete statement and, where possible, the version of the documentation.

For pricing, availability, legal obligations and product versions, the check should be repeated close to the time of publication. For academic definitions or standards, we track the edition and possible replacements. If the sources conflict, we don't automatically choose the newer page; we verify authority, scope and effective date.

An AI response may cite an accessible but outdated page. That's why "cited" and "correct on test date" are separate fields. Citation makes verification possible; it doesn't do it for us.

How we measure citations without inflating the result

The same response can display three URLs from the same domain. That's why we report separately:

  • citation occurrence share — domain citation occurrences ÷ all citation occurrences;
  • unique URL share — domain unique URLs ÷ all unique URLs;
  • domain presence — conversations with at least one domain citation ÷ eligible conversations;
  • claim support coverage — verified claims that have adequate support ÷ verified claims.

The denominators follow the guidelines in the AI KPIs and their denominators guide. We do not combine occurrences, URLs and conversations into a single "citation score".

Minimum source ledger

For each observation we keep:

  1. prompt, conversation and date;
  2. engine, interface, language and geography;
  3. exact claim analyzed;
  4. displayed URL and domain;
  5. position or citation mark, if any;
  6. source type;
  7. date of publication and date of access;
  8. match between source and assertion;
  9. status: supported, contradicted or unconfirmed;
  10. capture or export that allows rechecking.

The ledger documents what was observed. It must not claim to reveal internal weights, training data, or all pages processed by the model.

Conclusion

A brand can be discovered without being cited, cited without being named and mentioned without the source being identifiable from the response. These states must be measured separately.

A sound strategy does not aim only for “more citations”. It publishes accurate first-party sources, earns independent validation, keeps official information current and builds pages that clearly answer intermediate questions.

In an audit, we say only what we can prove: which URL was displayed, which claim it supported, in which conversation and on what date. The rest remains a hypothesis.

Sources

Last source check: 2 August 2026. The taxonomy is editorial methodology and does not assert that AYSA already exposes all states or fields described.