Methodology

How we measure - so a number means something.

Every figure Clear Cited publishes carries its method, its variance, the engine set, and a date. This is that method - the canonical reference every "how we measure" link points to. Our AI Search Optimization (AEO/GEO) work covers the five AI engines - plus Google's AI surfaces, measured separately (detailed below).

The short version. We run your real buyer prompts across the AI engines your buyers actually use, as many times as it takes to reach a stable estimate, and report the median with metric-matched 95% confidence intervals. Same prompts, same way, every month - so a change is a result, not a screenshot.

The engines

Five engines, measured the same way

The full Index and Full Audit measure five: ChatGPT, Perplexity, Claude, Gemini, and Grok - the assistants B2B and developer buyers actually use. Smaller offerings measure fewer by design (a free teardown covers three, a Starter audit two to three), so the engine count varies by offering - we never claim one universal number, and we never blend in engines we don't measure.

We weight what your buyers actually use. Grok currently drives near-zero B2B referrals, so we measure and report it for completeness but de-emphasise its weight in the headline picture.

ChatGPT
Perplexity
Claude
Gemini
Grok

Google AI surfaces

Google AI Overviews & AI Mode - measured separately

Google AI Overviews and AI Mode are the question we get asked about most - and they are a different surface, not one of the five assistants above. So we measure them on their own track, with their own sample size, 95% confidence interval and date, and never sum them into the five-engine share-of-model.

We report Google AI Overviews as its own surface, never blended into a cross-engine score — it only fires on some searches, and a click lands in your analytics as ordinary organic traffic, so we treat measured AI referral as a floor, not a ceiling.

A different surface

An AI Overview is a zero-click answer block on a Google results page - not a conversational assistant. Adding it into share-of-model would compare two different things, so we keep two clearly-labelled tracks and report each on its own terms.

A floor, not a ceiling

A click from an AI Overview reaches your analytics as google / organic, so referral tracking under-counts it - measured AI referral is a floor, not a ceiling. We pair the measured citation-presence with Google's first-party Generative-AI impressions (Search Console) - shown side by side, never blended into one number.

We measure 5 AI engines and report 2 answer surfaces — Google AI Overviews and Microsoft Copilot — beside them, never inside them. Most vendors publish one flat roster that mixes model families with product surfaces. We separate them: an engine is one foundation-model family we query directly, a surface is a product that renders an answer — often routed across several model families it does not disclose, and often correlated with plain search rank. Summing correlated surfaces into an engine count inflates coverage, so we never do it.

Both surfaces are positioned, not measured today, and we would rather say so than publish a number we were not entitled to take. For Microsoft Copilot the product we would measure is the PUBLIC consumer Microsoft Copilot (copilot.microsoft.com) - Bing-grounded, no sign-in, no tenant data. Never the tenant-scoped M365 Copilot / Work IQ. Most vendors that list it never say which product they query; we do — because the tenant-scoped Microsoft 365 version answers from the querying organisation’s own files, which would measure our documents rather than the market, so we refuse that source outright. The public product is covered by terms that prohibit automated access, so we do not query it either.

The procedure

Adaptive per-engine sampling, with metric-matched CIs

Why adaptive. Not every AI engine is equally consistent. Ask ChatGPT the same question ten times and you may get ten different lists; ask Perplexity and the answers barely move. So we don't run every engine a fixed number of times — we run each as many times as it takes to reach a stable estimate. More runs where answers vary, fewer where they're stable, never below a 5-run floor. That's why our numbers hold up when you re-check them.
External corroboration. This is not only our finding. In independent research published 28 January 2026, SparkToro (Rand Fishkin) with Gumshoe.ai (Patrick O’Donnell) had 600 volunteers run 12 different prompts through each of 3 tools a combined 2,961 times over November and December 2025, and reported “a <1 in 100 chance that ChatGPT or Google’s AI, if asked 100X, will give you the same list of brands in any two responses.” The same study found the consideration set is far steadier than the list: in a follow-up of 994 responses about headphones, the leading brands appeared 55–77% of the time. That pairing is the whole argument for measuring share of model rather than a rank — the ordering is noise, the frequency is signal. Read the study → (external research, not ours)

Real buyer prompts

We measure the category questions your buyers actually ask - not vanity terms. Audits surface the verbatim prompts.

Runs spent where they matter

AI answers are non-deterministic, so a single response is noise. A client audit runs at least 10 per engine (12 by default) and spends more runs on high-variance engines, never below a 5-run floor; the public Index runs a fixed 500 answers per category.

Metric-matched 95% CIs

Every presence figure carries a Wilson 95% confidence interval. Share of model carries a percentile bootstrap 95% confidence interval in the paid audit, and a Wilson 95% confidence interval in the public AI Visibility Index — the method matched to the metric and named per product, both computed from the actual runs.

Where our own intervals are too narrow. Share of model is a ratio over a pooled mention denominator whose observations are clustered — one answer may name several products, and repeated runs of a prompt are correlated. The paid audit already resamples at the answer level with a percentile bootstrap. The public AI Visibility Index still reports a Wilson interval, and Wilson assumes an independence that a clustered denominator does not have, so the Index's published share intervals are NARROWER than a cluster-corrected estimate would give. Bringing the Index onto the same clustered bootstrap the audit already uses is planned for a future measurement cycle; we have not restated the released figures, because a DOI points at them. We state this because it works against us: wider intervals would make MORE of our published rankings statistically indistinguishable, not fewer.
Non-determinism, stated plainly. The same prompt can yield different answers run to run. That's why we report a median and a confidence interval, never a one-off screenshot - and why we can't promise rankings, only measured, reproducible visibility. The decision rule: overlapping intervals are noise; tight, separated intervals are real movement.

The method, producing a live result

What the procedure actually outputs

The same steps, run on a live category — measured share-of-model, its confidence intervals, and the sample behind it.

Run on a real category, the procedure above produces this — the CI/CD leaderboard by share-of-model, with its 95% intervals, engine set, and date. Not a mockup; the live dataset, refreshed in place.

AI Visibility Index — CI/CD platforms · share-of-modelmeasured snapshot · 2026-08-07
#ProductShare of modelShare95% CI
1GitHub Actions21.5%19.7–23.3
2GitLab CI/CD20.0%18.3–21.8
3CircleCI17.8%16.2–19.5
4Jenkins16.1%14.5–17.7
5Argo CD7.3%6.3–8.6
6Buildkite6.4%5.5–7.6
Engines
5 of 5
Prompts
10
Answers
500
Runs/engine
10
Interval
Wilson score interval, 95% (z=1.96)

One of nine measured categories in the public AI Visibility Index. Point-in-time; engines change. · See the full Index →

What we report

Six measurement layers, one monthly report

Share-of-model is the headline, but visibility isn't one number. Every retainer report assembles these, each labelled measured vs modeled.

Share-of-model + rank

Your share of qualifying AI answers vs the named field, with rank, across the five engines.

Citation-source split

Where the citations come from, split owned / earned / community — each with its own CI.

Mention quality

Sentiment, prominence, and hallucination flags on each mention (modeled, labeled).

Off-site authority gap map

The high-citation off-site sources you're absent from, prioritized.

AI-referral attribution

Measured AI referral traffic with the attribution-gap disclosure (honest 'not connected' until GA4 is wired).

AI-bot / crawler analytics

Which AI crawlers reach your pages, and whether they're allowed to.

The model

The 7-KPI AEO model

"Visibility" isn't one number. We score seven, so you can see why share of model moves - presence, endorsement, citation and position each tell a different part of the story.

01

% citing your URL

How often an answer links a claim to a page you own.

02

% naming your brand

How often the answer names you at all in the category.

03

Share of model

Your share of qualifying answers vs the named field - the headline metric.

04

Sentiment

Whether you're recommended, listed neutrally, or cautioned against.

05

Presence quality

Mention depth + source quality + data richness behind the mention.

06

Brand recognition

Whether the engine describes you accurately and consistently.

07

Market position

Where you rank in the answer's shortlist against competitors.

Defensibility

Why the public Index is trustworthy

Pre-registered

The prompt set, the brand universe, and the metrics are fixed before each run, so results can't be cherry-picked after the fact.

No conflict of interest

Clear Cited excludes its own clients from the public ranked Index. We disclose the policy and never rank a brand we're paid to promote.

Comparability breaks, disclosed

If the prompt set or engine set changes between editions, a versioned note marks where two editions stop being directly comparable.

Check us the way you'd check anyone

None of this is worth anything if you only take our word for it. The 12 questions to ask anyone selling AI visibility - answered here for us, with the link to each answer, including the three where the honest answer is partly.

Open data, CC BY 4.0

Every leaderboard ships a CSV + JSON under CC BY 4.0; we never publish a figure for an engine that returned too little data, and every number is labelled measured vs modeled.

Honesty

What every figure carries

Definitions

What we mean by a citation

Three teams have split the word “citation” on three different axes, all under one word. That is a large part of why published estimates of the same nominal quantity differ by more than an order of magnitude.

WhoThe splitThe axis
AhrefsFound in → Cited inRetrieval stage vs generation stage
Semrush, with Kevin IndigCited (linked) vs mentioned (named)Linking vs naming, within the answer
EvertuneFoundational knowledge vs real-time retrievalParametric vs retrieved

The retrieved-versus-cited distinction is Ahrefs' and is shipped in their product. What is ours is the synthesis: that three different axes are in use under one word, which is why the published numbers disagree.

Which one we measure

We measure the Cited (linked) vs mentioned (named) sense. We measure the second sense — named as a source in the returned answer, as extracted by our pipeline. We do not measure retrieval into context. That is a scope choice, and we would rather say so plainly than dress it up as an impossibility: when we last checked the official API documentation, three of the five engines we measure exposed the retrieved set separately from the cited set, and the other two did not document it either way. So part of this is measurable and we have not measured it.

Retrieval into context is not observable for most of this roster: 3 of the 5 engines we measure expose the retrieved set separately from the cited set; 2 do not document it either way, so we record them as not verified rather than assuming an answer. Read from each vendor’s own API documentation on 2026-08-02.

EngineRetrieved set observable?Source
ChatGPTVERIFIED-EXPOSEDtheir docs
ClaudeVERIFIED-EXPOSEDtheir docs
GrokVERIFIED-EXPOSEDtheir docs
GeminiUNVERIFIEDtheir docs
PerplexityUNVERIFIEDtheir docs

Open questions

What we haven't measured

These are open, and they stay listed until they close. We would rather you found them here than found them yourself.

The correction we made to our own work. Our own citation-extraction layer failed unevenly across two of the five engines we measure: one returned zero distinct cited domains in all nine categories, and a second recovered only one or two against a list cap of five. We found it while auditing our own measurement pipeline, before the paper had been submitted anywhere, and withdrew the two citation-level findings that rested on it — then reported the extraction failure itself as a result, because citation-share figures are published across this sector with no disclosure of extraction quality. The share-of-model findings come from brand naming rather than citations; they were unaffected and are unchanged.

How often is a source retrieved into a model's context but never named in the answer?

We do not measure retrieval at all — we measure what the answer names. This is a scope choice and not a limit of the engines: when we last read the official API docs, three of the five engines we measure exposed the retrieved set separately from the cited set, and two did not document it either way.

How much wider would the public Index's share-of-model intervals be if they were cluster-corrected?

The Index computes a Wilson interval over a pooled mention denominator whose observations are clustered, so its published share intervals are narrower than a cluster-corrected estimate would give. The paid audit already resamples at the answer level; bringing the Index onto the same estimator is a future measurement cycle, and we have not restated the released figures because a DOI points at them.

What share of answers do the two answer SURFACES carry, on their own terms?

The two answer surfaces in our roster are positioned, not measured. A surface renders an answer but is not one foundation-model family, so it is never summed into an engine count, and we publish no share-of-model figure for either. They are named, and the reason they are kept apart is explained, in the separate-surface panel above.

Accuracy over time

When an engine's behaviour changes, we tell you

AI answer engines update their models continually, so a brand's visibility can shift without any change on your side. Because we measure on a fixed cadence with confidence intervals, we can tell a real behaviour change from ordinary run-to-run variation. When an engine's answers move materially beyond those intervals, we re-baseline the affected measurements and proactively tell the clients it affects — with the before-and-after we measured — so your reporting stays honest and comparable over time.

Measurement is half the system

Measuring is where it starts, not where it ends

This page is the measurement half. What we do with those numbers — attribute the citations each piece earns, learn what wins, refresh on cadence, and prove it back to you — is the full intelligence loop. See how we analyze →

See it run on your brand.

A free teardown applies this exact method to your category.

Get a free teardown

Keep reading

Most AEO optimizes only the slice of citations you own. Here's the whole map — all four tiers AI cites from, and how much of it we actually work.

Clear Cited works all four tiers

Coverage 100% · Control ~86% · Influence ~94% · Earned (best-effort)

Most AEO optimizes the ~44% you own. Clear Cited works all four tiers.

Control ~86% · influence ~94% · the last ~6% earned — worked, not guaranteed.

Source: Yext (6.8M citations), 2025-10 — a third-party citation-tier mix, not measured by our AI Visibility Index (which tracks share-of-model); illustrative of the tier split until we publish our own measured citation-source mix

Coverage is the scope of what we work — never a control or ranking guarantee. · See the full Citation Control Map →

See it measured — and take the report

The method above, running live — and the report you can cite.

The measurement deep dive — sampling, intervals, reproducibility. (2:37) · captioned film — no audio by design · illustrative data, labelled in-film
Watch an AI answer get measured, end to end. (0:45) · captioned film — no audio by design · illustrative data, labelled in-film

The State of AI Search — measured report (PDF)