Credible Roots / AI Search Visibility Metrics
Measurement
AI search visibility metrics and KPIs, defined properly
There is no rank inside ChatGPT, so every metric borrowed from SEO reporting is either meaningless or a lie by analogy. These are the measurements that do survive contact with a probabilistic system, each with the formula, the sample size it needs, and what it cannot tell you.
AI search visibility is measured against a fixed prompt set, not against keywords, because answer engines retrieve and compose rather than rank. The five measurements that hold up are: presence rate, the share of prompts in a fixed set where your brand is named at all; citation share, your links as a proportion of all sources cited across those answers; substitution set, the named competitors who appear when you do not; crawler reachability, a per-engine yes or no with the date it was checked; and corroboration count, the number of independent third-party sources stating the claim you want quoted. All of them require repeat sampling, because the same prompt returns different answers on different runs — a single run tells you almost nothing, and we treat five runs per prompt as the floor for a reportable figure. Anything presented as a ranking position inside an answer engine is fabricated, because no ranking exists to hold. The honest caveat on all of it: published research puts AI referral traffic at roughly 1% of total web traffic, so these are positioning metrics, not traffic metrics, and should be reported as directional evidence rather than as a number that must rise every month.
- There is no rank. Any "position in ChatGPT" figure is invented.
- Five runs per prompt is the floor. One run is an anecdote, not a measurement.
- Freeze the prompt set. A set edited between reports can show anything.
- Report who gets named instead of you. It is the most actionable number available.
On this page
Why SEO metrics do not transfer
A search engine produces an ordered list, so position is a real property of the result and rank tracking is a legitimate measurement. An answer engine retrieves passages at the moment of asking, composes a reply, and attributes a handful of sources. Nothing in that process produces an ordering you could occupy.
This has three consequences that most AI visibility reporting ignores.
- Position is undefined. There is no slot one to hold. Being mentioned in the second sentence rather than the fifth is worth noting, but it is a property of one generated answer, not a standing you keep.
- The result is not stable. Ask the same question twice and the wording, the sources and sometimes the named brands change. SEO metrics assume a result that persists between checks. These do not.
- The query is not a keyword. People ask answer engines long, conversational, context-carrying questions. A keyword list is the wrong unit; a prompt set is the right one.
So the measurement problem is closer to polling than to rank tracking: you sample a population of possible answers and report a proportion with a known sample size, rather than reading a position off a board.
The five metrics that hold up
These are the definitions we use in client reporting. They are ours rather than an industry standard, because no industry standard exists yet; every formula is stated so you can reproduce or dispute it.
1. Presence rate
Definition: the share of prompts in a fixed set where the brand is named anywhere in the answer.
Formula: prompts where the brand appears at least once, divided by total prompts in the set, across all runs. A set of 40 prompts run 5 times gives 200 observations; a brand named in 30 of them has a presence rate of 15%.
What it cannot tell you: whether the mention was favourable, or whether it was the recommendation rather than an aside. Presence is necessary, not sufficient.
2. Citation share
Definition: your domain as a proportion of all source links cited across the answers in the set.
Formula: citations to your domain divided by total citations returned across all runs. Count the distinct links, not the answers, because one answer may cite you twice.
Why it matters more than presence: being named is reputational; being cited is retrieval. A rising citation share means the engines are actually fetching and using your pages, which is the thing on-site work can move.
3. Substitution set
Definition: the named entities that appear in answers where you do not.
Formula: not a ratio. A frequency-ordered list of competitors, with the count of prompts each appeared in and the sources those answers cited.
Why this is the most useful number on the page: it converts an absence into a work list. If four of your five substitutes are cited via the same trade publication, you have learned exactly where the corroboration gap is.
4. Crawler reachability
Definition: per engine, whether that engine’s crawler can actually fetch your pages, with the date checked.
Formula: a yes or no per agent — GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, Applebot-Extended — established from server logs, not inferred from robots.txt.
The trap: robots.txt is a statement of intent, not evidence of access. A permissive file tells you nothing if a CDN bot manager or security product is returning 403 at the edge. The only confirmation is a request from that user agent appearing in your logs.
5. Corroboration count
Definition: the number of independent third-party sources stating the specific claim you want quoted.
Formula: a count of distinct domains, excluding your own properties, your press releases and anything you paid for or placed.
Why it is a leading indicator: published research consistently finds citations skewing toward earned third-party coverage rather than a brand’s own domain, which makes this the metric that moves first and the one most reporting omits entirely.
Variance: why one run tells you nothing
This is the part that separates measurement from screenshotting, and almost nobody reports it.
Answer engines are probabilistic. The same prompt, asked twice, can return different wording, different sources and sometimes a different set of named brands. A single run of a single prompt is one draw from a distribution. Reporting it as a result is like calling an election from one voter.
The practical consequences:
- Repeat every prompt. We treat five runs per prompt as the floor for a reportable figure, and more where the answer is visibly unstable between runs. Fewer than that and month-to-month movement is indistinguishable from noise.
- Report the sample size with the number. "Presence rate 15%" is not a finding. "Presence rate 15%, 40 prompts, 5 runs, 200 observations, collected 3–4 September" is.
- Clear the context between runs. Personalisation, chat history and memory features all contaminate a sample. Each run should start clean.
- Fix the model and record it. Results differ between models and between versions of the same model. A report that does not name what it asked is not reproducible.
- Expect movement without cause. A few points either way between months may be variance rather than progress. Treat a single month’s change as a signal to look, not as a result.
None of this makes the measurement worthless. It makes it a sampled estimate, which is an ordinary thing to report honestly and an easy thing to misreport confidently.
Building a prompt set that cannot be rigged
Every number above depends on the prompt set, which makes the prompt set the easiest thing in AI visibility reporting to quietly manipulate. Six rules make that harder.
- Write it before you measure. Prompts chosen after seeing which ones you win are not a measurement.
- Freeze it, and version it. Changes go in an appendix with dates. A set edited between reports can be made to show anything.
- Use the questions buyers ask, not the ones you want asked. "Who are the best X providers" is legitimate. "Why is [your brand] the best X provider" is not a prompt, it is a prompt for the answer you wanted.
- Include prompts you expect to lose. A set you win 90% of is measuring the wrong thing, and it will not show improvement because there is no headroom.
- Keep unbranded prompts in the majority. Branded prompts measure whether the engine knows you exist. Unbranded ones measure whether you are the answer, which is the thing being bought.
- Cover the whole decision. Definition questions, comparison questions, pricing questions and "is X worth it" questions all retrieve differently.
Forty to sixty prompts is usually enough to be stable without becoming unaffordable to re-run honestly at five runs each.
What belongs in the monthly report
| Report | With it, always |
|---|---|
| Presence rate | Prompt count, runs per prompt, collection dates, models queried |
| Citation share | The competing domains cited, by name |
| Substitution set | Frequency per competitor, and the sources their answers leaned on |
| Crawler reachability | Per agent, from server logs, with the date and the evidence |
| Corroboration count | The links, so the count can be audited |
| Changes shipped | Which pages changed, and what the passage said before |
| Prompt set version | Any additions or removals since last month, with the reason |
If a report contains a number without its sample size, the number is decorative.
Metrics to refuse
- "Ranking" or "position" in an answer engine. No ordering exists. Any figure here was constructed.
- A single screenshot of a good answer. One draw from a distribution, presented as a trend.
- "AI visibility score" with no published method. If the formula is proprietary, it is not a measurement, it is a brand asserting a number about itself.
- Sentiment scores with no stated classifier. Same problem, with an extra layer of modelling to hide in.
- Share of voice against an unnamed competitor set. The denominator decides the answer, so the denominator has to be published.
- Projected AI referral traffic. Published figures put AI referrals at roughly 1% of total web traffic. Forecasting your slice of that is guessing with a spreadsheet.
Questions people actually ask
What are the main AI search visibility KPIs?
Five hold up under scrutiny: presence rate, the share of a fixed prompt set where you are named; citation share, your links as a proportion of all sources cited; the substitution set, the competitors named when you are not; crawler reachability, a per-engine yes or no taken from server logs; and corroboration count, the independent third-party sources stating the claim you want quoted. Each needs its sample size reported alongside it.
Can you track a ranking position in ChatGPT?
No. Answer engines retrieve passages and compose a reply rather than ordering a list, so there is no position to occupy or track. Any tool reporting a rank inside an answer engine has constructed that number from something else and relabelled it.
How many times should each prompt be run?
Five is our floor for a reportable figure, and more where answers are visibly unstable between runs. These systems are probabilistic, so a single run is one observation rather than a measurement, and month-to-month movement from small samples is usually noise.
How big should a prompt set be?
Forty to sixty prompts is usually enough to be stable while remaining affordable to re-run honestly at five runs each. What matters more than size is that it was written before measuring, frozen between reports, majority unbranded, and includes prompts you expect to lose.
Should we measure AI visibility as a traffic channel?
Not yet. Published research puts AI referral traffic at roughly 1% of total web traffic. Measured as a traffic source the numbers will be small and disappointing; measured as position — whether you are the name that comes back when a buyer asks — the same work looks like what it is. Report it as the latter.
Do we need a paid tool for this?
No, though tools save time at scale. The measurements above can be collected by hand with a spreadsheet, a frozen prompt set and the discipline to clear context between runs. What a tool cannot supply is the honesty of the prompt set, which is where most of the error lives.
Related reading
- Answer engine optimization: the work these metrics measure
- How answer engines pick sources
- Which AI crawlers should you allow
- AI search visibility as a service
- Measuring distribution when attribution fails
- What published method looks like on a live client site
- Which agencies publish their method, and which do not
- Industry figures, traced to their primary sources
Sources
- OpenAI. Overview of OpenAI crawlers: GPTBot, OAI-SearchBot and ChatGPT-User.
- Anthropic. Crawler documentation and how site owners can block it.
- Google. Overview of Google crawlers and user-triggered fetchers, including Google-Extended.
- Credible Roots. The published citation research, with each figure attributed to its publisher. The roughly 1% AI referral traffic figure and the earned-media citation findings are set out there with sources.
How this was made: the metric definitions are ours and are published so they can be reproduced or disputed; they are not an industry standard and we do not present them as one. Crawler names are taken from each vendor’s own documentation, linked above. Drafted with AI assistance, then reviewed and approved by the named author before publication. Our editorial standards.
Want this run on your category?
We will build the prompt set with you, run it clean, and show you the presence rate, the citation share and — usually the uncomfortable part — exactly who gets named instead of you.
Book a 30 minute call