% citing your URL
How often an answer links a claim to a page you own.
Methodology
Every figure Clear Cited publishes carries its method, its variance, the engine set, and a date. This is that method - the canonical reference every "how we measure" link points to. Our AI Search Optimization (AEO/GEO) work covers the five AI engines - plus Google's AI surfaces, measured separately (detailed below).
The engines
The full Index and Full Audit measure five: ChatGPT, Perplexity, Claude, Gemini, and Grok - the assistants B2B and developer buyers actually use. Smaller offerings measure fewer by design (a free teardown covers three, a Starter audit two to three), so the engine count varies by offering - we never claim one universal number, and we never blend in engines we don't measure.
We weight what your buyers actually use. Grok currently drives near-zero B2B referrals, so we measure and report it for completeness but de-emphasise its weight in the headline picture.
Google AI surfaces
Google AI Overviews and AI Mode are the question we get asked about most - and they are a different surface, not one of the five assistants above. So we measure them on their own track, with their own sample size, 95% confidence interval and date, and never sum them into the five-engine share-of-model.
We report Google AI Overviews as its own surface, never blended into a cross-engine score — it only fires on some searches, and a click lands in your analytics as ordinary organic traffic, so we treat measured AI referral as a floor, not a ceiling.
An AI Overview is a zero-click answer block on a Google results page - not a conversational assistant. Adding it into share-of-model would compare two different things, so we keep two clearly-labelled tracks and report each on its own terms.
A click from an AI Overview reaches your analytics as google / organic, so referral tracking under-counts it - measured AI referral is a floor, not a ceiling. We pair the measured citation-presence with Google's first-party Generative-AI impressions (Search Console) - shown side by side, never blended into one number.
We measure 5 AI engines and report 2 answer surfaces — Google AI Overviews and Microsoft Copilot — beside them, never inside them. Most vendors publish one flat roster that mixes model families with product surfaces. We separate them: an engine is one foundation-model family we query directly, a surface is a product that renders an answer — often routed across several model families it does not disclose, and often correlated with plain search rank. Summing correlated surfaces into an engine count inflates coverage, so we never do it.
Both surfaces are positioned, not measured today, and we would rather say so than publish a number we were not entitled to take. For Microsoft Copilot the product we would measure is the PUBLIC consumer Microsoft Copilot (copilot.microsoft.com) - Bing-grounded, no sign-in, no tenant data. Never the tenant-scoped M365 Copilot / Work IQ. Most vendors that list it never say which product they query; we do — because the tenant-scoped Microsoft 365 version answers from the querying organisation’s own files, which would measure our documents rather than the market, so we refuse that source outright. The public product is covered by terms that prohibit automated access, so we do not query it either.
The procedure
We measure the category questions your buyers actually ask - not vanity terms. Audits surface the verbatim prompts.
AI answers are non-deterministic, so a single response is noise. A client audit runs at least 10 per engine (12 by default) and spends more runs on high-variance engines, never below a 5-run floor; the public Index runs a fixed 500 answers per category.
Every presence figure carries a Wilson 95% confidence interval. Share of model carries a percentile bootstrap 95% confidence interval in the paid audit, and a Wilson 95% confidence interval in the public AI Visibility Index — the method matched to the metric and named per product, both computed from the actual runs.
The method, producing a live result
The same steps, run on a live category — measured share-of-model, its confidence intervals, and the sample behind it.
Run on a real category, the procedure above produces this — the CI/CD leaderboard by share-of-model, with its 95% intervals, engine set, and date. Not a mockup; the live dataset, refreshed in place.
| # | Product | Share of model | Share | 95% CI |
|---|---|---|---|---|
| 1 | GitHub Actions | 21.5% | 19.7–23.3 | |
| 2 | GitLab CI/CD | 20.0% | 18.3–21.8 | |
| 3 | CircleCI | 17.8% | 16.2–19.5 | |
| 4 | Jenkins | 16.1% | 14.5–17.7 | |
| 5 | Argo CD | 7.3% | 6.3–8.6 | |
| 6 | Buildkite | 6.4% | 5.5–7.6 |
One of nine measured categories in the public AI Visibility Index. Point-in-time; engines change. · See the full Index →
What we report
Share-of-model is the headline, but visibility isn't one number. Every retainer report assembles these, each labelled measured vs modeled.
Your share of qualifying AI answers vs the named field, with rank, across the five engines.
Where the citations come from, split owned / earned / community — each with its own CI.
Sentiment, prominence, and hallucination flags on each mention (modeled, labeled).
The high-citation off-site sources you're absent from, prioritized.
Measured AI referral traffic with the attribution-gap disclosure (honest 'not connected' until GA4 is wired).
Which AI crawlers reach your pages, and whether they're allowed to.
The model
"Visibility" isn't one number. We score seven, so you can see why share of model moves - presence, endorsement, citation and position each tell a different part of the story.
How often an answer links a claim to a page you own.
How often the answer names you at all in the category.
Your share of qualifying answers vs the named field - the headline metric.
Whether you're recommended, listed neutrally, or cautioned against.
Mention depth + source quality + data richness behind the mention.
Whether the engine describes you accurately and consistently.
Where you rank in the answer's shortlist against competitors.
Defensibility
The prompt set, the brand universe, and the metrics are fixed before each run, so results can't be cherry-picked after the fact.
Clear Cited excludes its own clients from the public ranked Index. We disclose the policy and never rank a brand we're paid to promote.
If the prompt set or engine set changes between editions, a versioned note marks where two editions stop being directly comparable.
None of this is worth anything if you only take our word for it. The 12 questions to ask anyone selling AI visibility - answered here for us, with the link to each answer, including the three where the honest answer is partly.
Every leaderboard ships a CSV + JSON under CC BY 4.0; we never publish a figure for an engine that returned too little data, and every number is labelled measured vs modeled.
Honesty
Definitions
Three teams have split the word “citation” on three different axes, all under one word. That is a large part of why published estimates of the same nominal quantity differ by more than an order of magnitude.
| Who | The split | The axis |
|---|---|---|
| Ahrefs | Found in → Cited in | Retrieval stage vs generation stage |
| Semrush, with Kevin Indig | Cited (linked) vs mentioned (named) | Linking vs naming, within the answer |
| Evertune | Foundational knowledge vs real-time retrieval | Parametric vs retrieved |
The retrieved-versus-cited distinction is Ahrefs' and is shipped in their product. What is ours is the synthesis: that three different axes are in use under one word, which is why the published numbers disagree.
We measure the Cited (linked) vs mentioned (named) sense. We measure the second sense — named as a source in the returned answer, as extracted by our pipeline. We do not measure retrieval into context. That is a scope choice, and we would rather say so plainly than dress it up as an impossibility: when we last checked the official API documentation, three of the five engines we measure exposed the retrieved set separately from the cited set, and the other two did not document it either way. So part of this is measurable and we have not measured it.
Retrieval into context is not observable for most of this roster: 3 of the 5 engines we measure expose the retrieved set separately from the cited set; 2 do not document it either way, so we record them as not verified rather than assuming an answer. Read from each vendor’s own API documentation on 2026-08-02.
| Engine | Retrieved set observable? | Source |
|---|---|---|
| ChatGPT | VERIFIED-EXPOSED | their docs |
| Claude | VERIFIED-EXPOSED | their docs |
| Grok | VERIFIED-EXPOSED | their docs |
| Gemini | UNVERIFIED | their docs |
| Perplexity | UNVERIFIED | their docs |
Open questions
These are open, and they stay listed until they close. We would rather you found them here than found them yourself.
How often is a source retrieved into a model's context but never named in the answer?
We do not measure retrieval at all — we measure what the answer names. This is a scope choice and not a limit of the engines: when we last read the official API docs, three of the five engines we measure exposed the retrieved set separately from the cited set, and two did not document it either way.
How much wider would the public Index's share-of-model intervals be if they were cluster-corrected?
The Index computes a Wilson interval over a pooled mention denominator whose observations are clustered, so its published share intervals are narrower than a cluster-corrected estimate would give. The paid audit already resamples at the answer level; bringing the Index onto the same estimator is a future measurement cycle, and we have not restated the released figures because a DOI points at them.
What share of answers do the two answer SURFACES carry, on their own terms?
The two answer surfaces in our roster are positioned, not measured. A surface renders an answer but is not one foundation-model family, so it is never summed into an engine count, and we publish no share-of-model figure for either. They are named, and the reason they are kept apart is explained, in the separate-surface panel above.
Accuracy over time
AI answer engines update their models continually, so a brand's visibility can shift without any change on your side. Because we measure on a fixed cadence with confidence intervals, we can tell a real behaviour change from ordinary run-to-run variation. When an engine's answers move materially beyond those intervals, we re-baseline the affected measurements and proactively tell the clients it affects — with the before-and-after we measured — so your reporting stays honest and comparable over time.
Measurement is half the system
This page is the measurement half. What we do with those numbers — attribute the citations each piece earns, learn what wins, refresh on cadence, and prove it back to you — is the full intelligence loop. See how we analyze →
A free teardown applies this exact method to your category.
Get a free teardownKeep reading
Most AEO optimizes only the slice of citations you own. Here's the whole map — all four tiers AI cites from, and how much of it we actually work.
Clear Cited works all four tiers
Coverage 100% · Control ~86% · Influence ~94% · Earned (best-effort)
Most AEO optimizes the ~44% you own. Clear Cited works all four tiers.
Control ~86% · influence ~94% · the last ~6% earned — worked, not guaranteed.
Source: Yext (6.8M citations), 2025-10 — a third-party citation-tier mix, not measured by our AI Visibility Index (which tracks share-of-model); illustrative of the tier split until we publish our own measured citation-source mix
Coverage is the scope of what we work — never a control or ranking guarantee. · See the full Citation Control Map →
The method above, running live — and the report you can cite.