Blog · 2026-06-14 · By Logan Adams, Founder · 4 min read
Why one ChatGPT screenshot doesn't prove your AI visibility
The short version: a screenshot of one AI answer is a single draw from a distribution, not evidence. Ask the same question again and the names, order, and sources can change. To know whether AI actually recommends you, you have to measure it like a statistic — repeat the prompt, report the median with a confidence interval, do it per engine, and re-check on a cadence. Here's why, and what "measured properly" looks like.
Someone shows you a screenshot: they asked ChatGPT for the best tool in your category, and your competitor came up instead of you. Or the reverse — you came up, and it feels like a win. Either way, that screenshot proves almost nothing.

AI answers are non-deterministic
Large language models sample their output. Ask the same question twice and you can get two different lists, in a different order, citing different sources. A single answer is one draw from a distribution; treating it as fact is like calling a coin biased after one flip.
It gets less stable, not more, in the real world. The answer can shift with the exact wording of the prompt, the user's location and history, the time of day, which model version is serving, and which pages the engine happened to retrieve that minute. Two honest people can screenshot the same question an hour apart and "prove" opposite things.
A tale of two screenshots
Picture a founder and their competitor both screenshotting "best [category] tool" on the same afternoon. The founder's capture shows them in second place; the competitor's shows them absent entirely. Both are real. Both are useless on their own — because neither tells you the rate: out of a hundred honest asks, how often does each name actually appear? That rate, per engine, is the thing worth managing. A screenshot is a single sample of it with the error bars cropped off.
What measurement actually requires
To say anything real about whether AI recommends you, you have to treat it like the statistical question it is:
- Repeat each prompt. We run every buyer question 10 or more times per engine. One run is an anecdote; a distribution is data.
- Report the middle, not the lucky run. We report the figure with a confidence interval named per product — a percentile bootstrap 95% interval for share-of-model in the paid audit, a Wilson 95% interval for share in the public Index, and a Wilson 95% interval for presence in both — so you see the range, not a cherry-picked best or worst case.
- Measure per engine. ChatGPT, Perplexity, Claude, Gemini, and Grok pull from different sources and give different answers; strong visibility on one says little about another. We've measured exactly this — the leader changes with the engine.
- Use real buyer prompts. Measure the questions your buyers actually ask ("best X for a small team," "X alternative"), not vanity queries that contain your brand name.
- Re-measure on a cadence. AI answers drift as the models and their sources change — a substantial share of the sources AI cites turns over from one month to the next (Profound) — so a number from last month may already be stale.
Why this matters for the work, not just the report
The discipline isn't academic. If you can't measure reliably, you can't tell whether anything you did actually moved your visibility — you're just guessing and hoping. And because the inputs are mostly off-site (roughly 95% of the citations behind AI answers come from third-party pages, not your own marketing — Otterly), the temptation to judge progress from a flattering screenshot is strong and misleading.
Reproducible measurement is what turns AI visibility from vibes into something you can manage: a baseline, a target, and evidence that the work is paying off. It's also what keeps everyone honest — a median with a confidence interval can't be cherry-picked the way a single screenshot can.
That's the whole basis of how we operate. No single screenshots, no black-box scores — just a number you can trust, measured the same way every time, against your named competitors.
FAQ
How many runs is "enough"? Enough that the confidence interval is tight enough to act on. We start at 10+ runs per prompt per engine and widen for noisier categories; the CI tells you when you have signal rather than a lucky draw.
Isn't a screenshot fine as a quick gut-check? As a prompt to go measure, sure. As proof of where you stand — or as a before/after for work you're paying for — no. It has no rate and no error bar.
Why measure competitors too? "Share of model" is relative: your visibility only means something against the other names in the category. Measuring the field is how you know whether you're leading, mid-pack, or invisible on each engine.
See it on your own domain
Want your real AI visibility, measured properly across the engines your buyers use? Get a free teardown and we'll show you where you stand against your named competitors — reproducibly, not from a single screenshot.
Last updated: 2026-06-24 (published 2026-06-14). We re-check figures on a cadence because AI engines change continuously.

Comments
No comments yet — be the first, below.