Mohammed TetoAI Reputation & Authority
Research methodology

A result is published with everything needed to check it.

This page is the standard that any measurement published on this site has to meet. It is written before any result, so that a result cannot shape the rule.

This methodology sets the conditions under which an observation of AI-generated answers may be published on this site: the question is stated first, the prompts are frozen, every run is recorded with its conditions, the coding rules are written down, and results are reported per engine with their sample size. It applies to public research. Client work is confidential and is not published here. As of October 4, 2026, no measurement has been made or published under this protocol.

By · Published · Updated

Why a protocol is needed

Generated answers vary from one run to the next, differ by engine, and depend on how the question is worded. A single screenshot proves that one answer was given once. The reasons are set out in why rank tracking does not transfer. A protocol does not remove that variation. It makes it visible, and it lets someone else repeat the measurement and compare.

The ten rules

  1. The question comes first. What is being measured, about which entities, is written down and dated before any prompt is run.
  2. The prompt set is frozen. Prompts are fixed, numbered and versioned. A change of wording is a new version, and results from two versions are not merged. The families are described in the prompt library.
  3. Each engine is measured separately. Results are never blended across engines.
  4. Every run is recorded with its conditions. Prompt, engine, interface (application or API), whether the session was signed in, model version where it is shown, date and time, language, location, whether web search was active, the full answer and the sources cited. The fields are listed in what is recorded.
  5. The number of runs is fixed in advance. It is set before the first run, is the same for every prompt and every engine, and is stated with the result.
  6. Coding rules are written before coding. They are published with the result, as set out below.
  7. The full answers are kept. They are made available with the result, apart from passages about people other than the subject. An observation can be re-coded from an answer; an answer cannot be recovered from a table.
  8. Results state their sample. Every proportion is given with the number of prompts and runs behind it.
  9. Limits are part of the result. Each publication says what the measurement cannot show.
  10. Errors are corrected in the open. A correction is dated and entered in the change log.

How an answer is coded

Coding turns an answer into observations. The rules are defined so that two people can record an answer the same way; that agreement has not been tested yet, and a first publication would have to report it.

Observation Rule Values
Presence Is the entity named in the answer, under its name or a listed variant? Yes or no, per run
Identity When named, is it the right entity, and not a namesake? Right, wrong, mixed
Accuracy Each fact stated is compared with the reference version: name, role or activity, organization, dates Correct, incorrect, outdated, unsupported, per fact
Sources Which domains are cited, classified by who controls them The seven kinds listed in kinds of source
Others named Which other organizations appear in the same answer A list, per run
Position When the entity is named, where does its first mention fall in the answer? First named, among several named, in passing; reported only as a distribution over the fixed number of runs, never from one run

The reference version is the set of facts the entity can document, with the accepted variants of its name. It is fixed before coding. "Outdated" means a fact that was once true is given as current. "Unsupported" means the answer states something that the reference version neither confirms nor contradicts; it is counted, because an invented claim is the most serious error an answer can contain.

An answer coded "wrong" for identity counts as present for the name and is excluded from the accuracy count, which would otherwise measure someone else.

Two things are not coded. Tone and sentiment: no rule has been written here that two readers would apply to them the same way. Consistency across engines: it is read from the per-engine tables and is not given a value of its own.

How results are reported

As proportions of runs. The number of runs in which the entity is named, out of the number of runs made, and not a rank. Presence, wherever this site reports it, is the share of runs, across the prompt set, in which the entity is named.

Per engine. One table per engine, with its date range.

With the prompt set attached. The prompts, their version and the coding rules are published with the result, so that the measurement can be repeated.

Without a blended score. No single index is published. A composite number hides which observation moved, and would need a statistical validation that has not been done.

With the sample in view. A share-of-voice figure is not produced from fewer than 50 prompts, the minimum one provider gives; the thresholds are set out in what a valid reading requires.

What is never published

  • Client data without written consent Named or anonymized.
  • A measurement about a named person without that person's written agreement Other people who appear in answers are not published by name.
  • Unverified cases No result is attributed to this practice unless each claim in it has been verified.
  • A result without its method If the prompts cannot be published, the result is not.

Limits of the method itself

The prompts are a model. They represent questions people may ask. They are not a record of the questions asked.

Conditions are only partly controlled. Engines do not always show which model answered, and an answer can depend on factors the observer cannot see.

Citations are the visible part. An answer can draw on a source without naming it; the distinction is explained in a source used and a source cited.

A measurement dates quickly. Engines and sources change. A result is a statement about its date range.

The author has an interest. These measurements would be published by an advisor who sells work in this field. Publishing the prompts, the rules and the sample is the safeguard: the reader does not have to take a result on trust.

Where it sits

This page governs what is published in Research. The terms it uses are defined in the glossary. The method it applies is explained for a general reader in AI search intelligence, and belongs to the first and last layers of The Authority Architecture System.

Frequently asked questions

Has any measurement been made or published under this methodology?

No. As of October 4, 2026, no measurement has been made or published, and this page states the rules. Any result will be published with its prompts, conditions and coding rules.

Why is there no overall score?

A composite score hides which observation changed, and it would need statistical validation that has not been done. Results are reported observation by observation and engine by engine.

Why is sentiment not measured?

Because no rule has been written here that two readers would apply to tone the same way. The protocol keeps observations that can be defined precisely: presence, identity, accuracy, sources and others named.

Can someone else reproduce a result?

That is the purpose of publishing the prompts, the conditions and the rules. A repeat run will not return identical answers, since answers vary. Whether its proportions are comparable is what the published sample lets a reader judge.

Understand how search and AI currently interpret your authority.

Request a confidential assessment of your visibility, entity signals, narrative consistency and source authority.

Request a Private Assessment

Selective engagements for leaders and organizations with material reputational stakes.