Free original research
The measured state of AI search
By Logan Adams, Founder Reviewed & updated Measurement-first: figures are median share-of-model with 95% confidence intervals. How we measure.
We publish what ChatGPT, Perplexity, Claude, Gemini, and Grok actually recommend — ranked by share of model, measured reproducibly across ChatGPT, Perplexity, Claude, Gemini, and Grok (adaptive per-engine sampling, at least 10 runs per engine with a 5-run floor; share-of-model carries a Wilson 95% CI over the pooled mention denominator, presence a Wilson 95% CI), with the data and method open.
Reports & datasets
State of AI Search
The headline finding: engines disagree on the “best tool,” and ~95% of the citations behind their answers are third-party. What that means for AI visibility.
LeaderboardsAI Visibility Index
Per-category share-of-model leaderboards (CI/CD, observability, and more), refreshed monthly. Datasets published CC BY 4.0.
Living studyDo directory listings change AI citation?
A pre-registered, living before/after measurement on ourselves — does buying directory listings actually change whether AI engines cite you? Null results published, method open to challenge, raw data ungated.
How we measure
Every buyer prompt runs 10+ times per engine; we report the median share of model with a 95% confidence interval. Consumer apps and APIs differ — we say so.
What we measure right now
Share-of-model leaders across the categories in our AI Visibility Index. Measured rows are medians over 10+ runs per engine; share-of-model carries a Wilson 95% confidence interval over the pooled mention denominator (presence a Wilson 95% CI). Illustrative previews are clearly marked and are not live measurements. Coverage: 9 measured, 0 illustrative.
| Category | Leader | Share-of-model | 95% CI | Status |
|---|---|---|---|---|
| AI observability tools | Datadog | 16.4% | 15.0–17.9 | Measured 2026-07-01 |
| API platforms | Kong | 20.3% | 18.6–22.1 | Measured 2026-07-01 |
| CI/CD platforms | GitHub Actions | 21.8% | 20.1–23.7 | Measured 2026-07-01 |
| CRM software | HubSpot | 22.3% | 20.6–24.1 | Measured 2026-07-01 |
| Databases | PostgreSQL | 22.3% | 20.5–24.1 | Measured 2026-07-01 |
| Feature flag platforms | LaunchDarkly | 17.2% | 15.7–18.8 | Measured 2026-07-01 |
| Incident management platforms | PagerDuty | 25.0% | 23.1–27.1 | Measured 2026-07-01 |
| Product analytics platforms | Mixpanel | 23.2% | 21.4–25.1 | Measured 2026-07-01 |
| Vector databases | Weaviate | 19.1% | 17.6–20.8 | Measured 2026-07-01 |
Browse every category and download the datasets (CC BY 4.0) on the AI Visibility Index. AI answers are non-deterministic; figures reflect a point in time and will drift. We do not guarantee any ranking or citation.
The research programme
A research programme, not a blog: every study is pre-registered before we look, ships with a limitations block, and is designed to be repeated — the single biggest structural gap in a landscape of one-off vendor snapshots. Publication doesn't depend on the direction of the result; a null publishes with the same production values as a win.
| Study | Tier | Status | Uniqueness | Evidences |
|---|---|---|---|---|
| AI Answer Volatility Index | Tier 1 | running | Replication Rigor Upgrade | measured, growth_retainer… |
| Cross-Engine Disagreement Index | Tier 1 | proposed | Replication Rigor Upgrade | measured, full_audit… |
| Newsletter/email presence vs AI citation | Tier 1 | proposed | Unique | growth_retainer, scale_retainer… |
| Schema markup on NEVER-cited pages | Tier 1 | proposed | Unique | growth_retainer, addon_technical_seo |
| Content-freshness refresh test (causal, matched control) | Tier 1 | proposed | Unique | addon_blog_post, growth_retainer… |
| Developer-doc / technical-content citation study | Tier 1 | proposed | Unique | growth_retainer, scale_retainer… |
| GEO industry citation-hygiene audit | Tier 1 | proposed | Unique | house authority |
| The directory study (does buying listings change AI citation?) | Tier 1 | running | Unique | addon_directory_boost, addon_directory_growth… |
| The negative-results ledger (what didn't work) | Tier 1 | running | Unique | measured, growth_retainer… |
| AI citation -> funnel/revenue | Tier 2 | proposed | Replication Rigor Upgrade | measured, growth_retainer… |
| Own-brand Reddit test (disclosed participation only) | Tier 2 | proposed | Unique | growth_retainer, scale_retainer… |
| Directory causal test across clients | Tier 2 | proposed | Replication | addon_directory_boost, addon_directory_growth… |
| PR placement -> citation lift | Tier 2 | proposed | Unique | addon_digital_pr, scale_retainer… |
| Brand mentions vs backlinks, cleanly separated | Tier 2 | proposed | Replication Rigor Upgrade | addon_backlink_domain, addon_digital_pr |
| Multi-baseline SCED across >=3 clients | Tier 2 | proposed | Unique | growth_retainer, scale_retainer… |
| State of AI Visibility (the annual flagship) | Tier 3 | proposed | Unique | measured, growth_retainer… |
| Citation-source composition BY VERTICAL | Tier 3 | proposed | Replication Rigor Upgrade | measured, full_audit |
| The leaderboard halo effect | Tier 3 | proposed | Unique | house authority |
Each study names the SKU it is the evidence base for — see the Evidence section on the relevant pricing page — or is marked house authority (a brand/category asset, not sales copy).
Want your own category measured?
A free teardown runs your buyer prompts across every engine.
Get a free teardown