# An AI answer is not a ranking. It has to be measured differently.

> How presence in AI answers is measured: what to record, which metrics mean something, how many prompts a reading needs, and what the numbers cannot say.

- Page: https://mohammedteto.com/ai-search-intelligence/
- Language: en — French version: https://mohammedteto.com/fr/mesure-recherche-ia/
- Author: Mohammed Teto (https://mohammedteto.com/#person)
- Published: 2026-10-07
- Updated: 2026-10-04

A generated answer has no fixed position one, and the same question can return a different answer an hour later. AI search intelligence is the method for measuring presence, accuracy and sources under those conditions.

AI search intelligence is the measurement of how an entity appears in answers produced by AI systems: whether it is present, how it is described, which sources are cited and which competitors are named beside it. It does for generated answers what rank tracking does for search results, with different instruments. Its output is a dated, repeatable record, not a score.

<a id="why"></a>

## Why rank tracking does not transfer

Three properties of generated answers break the usual instruments.

**The query is rewritten.** Google states that AI Overviews and AI Mode may issue multiple related searches across subtopics and data sources to develop a response ([Google Search Central](https://developers.google.com/search/docs/appearance/ai-features)). Measured figures are given in [why rankings do not guarantee citations](https://mohammedteto.com/ai-visibility/#rankings-vs-citations). The question you track is not the search the system runs.

**The answer varies.** The same prompt can produce a different wording, and different sources, from one run to the next. A single run is an observation, not a measurement.

**Google counts AI features inside Web search.** Its documentation states that sites appearing in AI features are included in the overall search traffic in Search Console, in the Performance report, within the "Web" search type. Search Console covers Google Search only, not the other assistants.

<a id="record"></a>

## What is recorded

Every run is stored with its conditions. Without them, two readings cannot be compared.

| Field | Why it is kept |
|---|---|
| Prompt, word for word | A change of wording is a change of test |
| Engine and, where known, model version | Engines rely on different sources |
| Date and time | Answers change as sources and models change |
| Language and location | The same entity can be described differently from one market to another |
| Whether web search was active | An answer from memory and an answer from retrieval are two different things |
| The full answer | Scoring can be redone later; a summary cannot |
| The sources cited | They show where the description comes from |

<a id="metrics"></a>

## The metrics, and what each one means

| Metric | Definition | What it does not tell you |
|---|---|---|
| Presence | The share of relevant prompts, over repeated runs, in which the entity is mentioned | Whether the mention is accurate or favorable |
| Accuracy | Whether name, role, activity and facts match the reference version | Whether the entity is recommended |
| Position in the answer | Where the mention falls: first, among several, in passing | Why |
| Citations | Which domains the answer cites, and whether the entity's own pages are among them | Whether uncited sources were used |
| Share of voice | The entity's mentions as a proportion of all mentions in a category prompt set | Whether the description is accurate; nothing reliable on a small prompt set |
| Competitors named | Which other entities appear in the same answers | Whether they are true competitors |
| Consistency | Whether the description holds across engines, languages and runs | Which version is right |

Research gives one example of how visibility can be quantified. The paper that introduced the term GEO measures a source's "position-adjusted word count": the share of an answer's words tied to that source, weighted by how early they appear ([Aggarwal et al., KDD 2024](https://arxiv.org/abs/2311.09735)). It is a research metric, built for controlled experiments, and it treats a mention at the top of an answer and one at the bottom differently.

<a id="sample"></a>

## How many prompts make a reading

It depends on the question asked.

**Describing one entity** needs a small, deliberate set: who is this person, what does this company do, is it credible, how does it compare. This is a diagnostic. It shows how the entity is described today, and it is what a first assessment uses.

**Measuring share of voice in a category** needs volume. LLM Pulse recommends 100 to 200 prompts, with 50 as a minimum for a defensible reading ([LLM Pulse, 2026](https://llmpulse.ai/blog/measure-ai-share-of-voice/)). Below that, a percentage is an anecdote with a decimal.

In both cases the prompts are grouped into families (entity, category, comparison, use case) and each run is repeated. The protocol is set out in [how AI visibility is tested](https://mohammedteto.com/ai-visibility/#testing).

<a id="limits"></a>

## What the numbers cannot say

**They are a model of demand, not a record of it.** The prompt set is written by the analyst. It represents what people may ask; it is not a log of what they asked.

**They do not explain themselves.** A fall in presence can come from a model update, a source that changed, or a competitor that published. The cause is found by reading the answers and their sources, not the chart.

**They are not comparable across tools.** Two platforms that use different prompts, engines and schedules can report different figures for the same entity. A trend is meaningful inside one protocol.

**They do not measure reputation.** An entity can be present everywhere and described wrongly. What the answer says is the subject of [AI reputation](https://mohammedteto.com/ai-reputation/).

<a id="decisions"></a>

## From reading to decision

A measurement is useful when it points to a layer of work.

| Reading | Where the work is |
|---|---|
| Absent from answers | Sources and evidence: see [generative engine optimization](https://mohammedteto.com/generative-engine-optimization/) |
| Present, described inaccurately | The entity and its sources: see [entity optimization](https://mohammedteto.com/entity-optimization/) |
| Cited sources are weak or outdated | Corrections at the source, and stronger independent references |
| Own pages rank for the question but are not cited | Page structure: see [answer engine optimization](https://mohammedteto.com/answer-engine-optimization/) |
| Described well in one engine, poorly in another | The sources that engine relies on |
| Described well in one language, poorly in another | Sources in that language and market |

<a id="next"></a>

## Where to go next

This page describes the method; the conditions under which a measurement is published on this site are set in the [research methodology](https://mohammedteto.com/research/methodology/). The ongoing engagement that applies it to one entity, month by month, is [AI Reputation Intelligence](https://mohammedteto.com/ai-reputation-intelligence/), described on its own page. AI search intelligence is one of the disciplines described in [AI visibility](https://mohammedteto.com/ai-visibility/), and corresponds to the first and last layers of [The Authority Architecture System](https://mohammedteto.com/method/): [Perception Intelligence](https://mohammedteto.com/method/#perception-intelligence), then [Monitoring and Governance](https://mohammedteto.com/method/#monitoring-and-governance). Which of the three disciplines a reading points to is summarized in [SEO, AEO and GEO compared](https://mohammedteto.com/compare/seo-vs-aeo-vs-geo/#choose). Related analyses are collected in [Insights](https://mohammedteto.com/insights/).

<a id="faq"></a>

## Frequently asked questions

### What is AI search intelligence?

It is the measurement of how an entity appears in AI-generated answers: presence, accuracy of the description, sources cited, competitors named and consistency across engines, recorded with the conditions of each run so that readings can be compared over time.

### How do I track my visibility in ChatGPT?

With fixed prompts, recorded conditions and repeated runs. The protocol is described in [how AI visibility is tested](https://mohammedteto.com/ai-visibility/#testing). This page adds what to record and how to read each metric.

### Can Google Search Console show my AI Overview traffic?

It is counted, inside the total. Google states that sites appearing in AI features are included in overall search traffic in Search Console, in the Performance report within the "Web" search type.

### How many prompts do I need?

For a diagnostic of one entity, a small deliberate set. For share of voice across a category, far more: see [how many prompts make a reading](#sample).

### Why do I get a different answer each time?

Generated answers vary from one run to the next, and the sources retrieved can change. This is why each prompt is repeated and why results are reported as a share of runs, not as a single result.

### Is AI share of voice a reliable metric?

Only with enough prompts, a stable protocol and a category defined in advance. With a small set it is unstable, and it never says whether the description is accurate.

---

Understand how search and AI currently interpret your authority. Request a confidential assessment of your visibility, entity signals, narrative consistency and source authority.

[Request a Private Assessment](https://mohammedteto.com/private-assessment/)
