Buyer's checklist · vendor-neutral

12 questions to ask anyone selling AI visibility

Every one of these is answerable in a sentence by a provider who measures, and unanswerable by a provider who guesses. Take the list to whoever you are evaluating — including us. Our answer to each is below, with the link to where we already published it.

Take me to the 12 questions Read our full method

How to use this

Ask all twelve. You are not looking for perfect answers — you are looking for specific ones. “We run it a bunch of times” is not a sample size. “Our proprietary score” is not a formula. A provider who cannot say what their confidence interval is computed from does not have one.

Watch for the two failure modes. The first is vagueness: an answer that describes a category rather than a number. The second, harder to spot, is a real-sounding number with no denominator — a percentage with nothing under the line, a rank with no idea how many things were ranked, a “score” whose inputs are not named.

We answer each one here honestly, and three of them we answer with partly. That is deliberate. A checklist a vendor writes to make itself look good is worth nothing; the only version worth publishing is the one we can fail.

The questions

1. How many times do you run each prompt, per engine?

Why it matters: one AI answer is one sample from a distribution. A single screenshot tells you nothing you can act on, and nothing you could reproduce next Tuesday.

Our answer: a client audit runs each buyer prompt at least 10 times per engine — 12 by default — spending more runs on engines whose answers vary more, and never dropping below a 5-run floor. The public Index is deliberately uniform instead: a fixed 10 runs × 10 prompts × 5 engines per category. Method · the Index.

2. What is your confidence-interval method, and is it named per metric?

Why it matters: “95% CI” is not a method. Different estimators suit different metrics, and a provider who uses one label for everything has probably not thought about which.

Our answer: every presence figure carries a Wilson 95% interval. Share of model carries a percentile bootstrap 95% interval in the paid audit and a Wilson 95% interval in the public Index — matched to the metric, named per product, both computed from the actual runs. We also publish where our own intervals are too narrow: the Index's share denominator is clustered, and Wilson assumes an independence it does not have. The caveat, in our words · Wilson · bootstrap.

3. Can I see the raw evidence behind a number?

Why it matters: a dashboard is a claim. The evidence is the answers the engine actually returned.

Our answer: partly, and here is exactly where the line is. Every public leaderboard ships its full CSV and JSON under CC BY 4.0, so anyone can recompute our aggregates. Verbatim prompts and the answers behind them surface in a paid audit, to the client whose brand they concern — we do not publish the full answer transcripts for the public Index. One living study, on whether directory listings change AI citation, ships its raw data ungated. Download a leaderboard · the research programme.

4. Is the prompt list fixed in advance, and is the scoring formula published?

Why it matters: the prompt set is the score. Change which questions you ask and you change who wins, without changing a single measurement.

Our answer: fixed in advance and hash-committed; published as a definition rather than an equation. The prompt set and the brand universe are frozen — hashed — before every run, so there is no post-hoc cherry-picking. Share of model is defined publicly: a product's share of all recommendations across every prompt and engine, over a pooled mention denominator. The verbatim prompt text is shown in the audit, not on the public page. The pre-registration rule · how share of model is defined.

5. Are you querying the API or the consumer app — and do you say which?

Why it matters: they are different products. An API answer and the answer your buyer sees in the app can differ in content, in sources, and in whether there are citations at all.

Our answer: we treat them as different surfaces and say so. The clearest case is the one where we refuse: Microsoft Copilot's public product is covered by terms that prohibit automated access, so we do not query it — which is why it is positioned rather than measured. Where that line is drawn · how we measure.

6. What geography are the probes run from?

Why it matters: AI answers are localised. A leaderboard measured somewhere your buyers are not is a leaderboard about someone else's market.

Our answer: the AI Visibility Index runs are measured in the US/English region, and an audit's money-prompt measurement uses the same setup — so the data reflects your market rather than ours. Stated on the US page.

7. Do you count AI “engines” and AI “answer surfaces” separately?

Why it matters: this is the quietest way to inflate coverage. Two products rendering the same underlying model are not two independent measurements, and summing them makes a roster look bigger than the evidence.

Our answer: five engines measured — ChatGPT, Perplexity, Claude, Gemini, Grok — and two answer surfaces reported beside them, never inside them: Google's AI surfaces (AI Overviews & AI Mode) and Microsoft Copilot. Seven properties, never summed into an engine count. Engines vs surfaces.

We report Google AI Overviews as its own surface, never blended into a cross-engine score — it only fires on some searches, and a click lands in your analytics as ordinary organic traffic, so we treat measured AI referral as a floor, not a ceiling.

8. How do you tell model drift apart from the effect of the work you did?

Why it matters: engines update continuously. Without an answer here, every improvement is claimable and no regression is anybody's fault.

Our answer: partly solved, and we say which part. Measuring on a fixed cadence with intervals separates a real behaviour change from ordinary run-to-run variation; when an engine moves materially beyond those intervals we re-baseline and tell the affected clients, with the before-and-after. What that does not yet give us is causal isolation when an engine update and a content change land together — the matched-control freshness test that would is on the research programme as proposed, not run. Our drift policy · the programme, including what is only proposed.

9. Are your paying clients excluded from your public rankings?

Why it matters: a public leaderboard that can be entered by paying the leaderboard's author is an advertisement.

Our answer: Clear Cited excludes its own clients from the public ranked Index, and does not rank its own category at all. The policy is disclosed rather than hidden. The conflict-of-interest rule · on the Index itself.

10. Is the study pre-registered — fixed before anyone looks at the data?

Why it matters: anything decided after the numbers arrive can be decided to flatter them.

Our answer: the prompt set, the brand universe and the metrics are fixed before each run, and publication does not depend on the direction of the result — a null publishes with the same production values as a win. The research programme.

11. Is the underlying data openly licensed, downloadable, and citable?

Why it matters: data you cannot download is a claim you cannot check. Data with no persistent identifier is a claim that quietly changes.

Our answer: every leaderboard ships CSV and JSON under CC BY 4.0, free, no email wall, with every figure labelled measured or modelled. The dataset is deposited and carries a DOI, which is why we do not silently restate released figures — something is pointing at them. The downloads · the dataset DOI.

12. Is your pricing public and fixed?

Why it matters: pricing that requires a call is pricing that varies by how much they think you will pay. It also tells you the measurement is not the product.

Our answer: every price is fixed and public, you buy online, and retainers are cancel-anytime with no lock-in beyond a prepay term you choose. One rung — genuinely non-standard multi-brand work — is contact-only, and we say so rather than hiding a number behind it. Pricing.

We asked these of the market, too

We checked what the providers in our own comparison set publish about their method, and scored ourselves on the same axes. That comparison is dated, sourced, and links to what we checked — so you can re-run it rather than take our word. See the disclosure comparison.

We do not characterise any provider's honesty here, and we have not named one on this page. The point of a buyer's checklist is that you use it, not that we grade the class.

Ask us all twelve

The fastest way to check the answers is to make us produce one. A free teardown runs your own buyer prompts across the engines and comes back with the numbers, the method and the date attached.

Get a free teardown