Convergence Blog September 2026
AI Search Visibility

AI Citation Measurement Crisis: Why Your AI Visibility Score Is Noise & How to Build a Proxy Stack That Survives It

Your AI visibility score is mostly noise. I ran 40 tests to find which AI search metrics hold still, and the proxy stack I use to report them.

¶ By Lillian Pierson, P.E. 12-minute read September 17, 2026 Page 01
Contents
The Convergence

Founder-tested growth plays, weekly.

Subscribe Free
Note: This content may be sponsored or contain affiliate links. If you purchase something after clicking an affiliate link, I may earn a small commission. Thank you for supporting small business.

5 assertions about your AI visibility score and AI visibility measurement, the data behind each one, and the exact calls to test them against your own domain.

On September 2 my AI visibility score read 12. Four days later, same tool and same 25 prompts, it read 22.

I published zero articles in between. So I pulled 40 consecutive runs spanning 5 months to find out what was actually happening.

I’m Lillian Pierson, a fractional CMO and licensed Professional Engineer who’s trained more than 2 million people in AI and data, and I’ve spent the past year tracking which sources AI engines actually cite.

What follows is what that data asserts, why it changes the way you should report AI search to your own stakeholders, and how to check every claim here against your domain in about 10 minutes.

The core takeaways of this article:

  1. Sampled vendor scores are more than likely just noise. Single-point AI visibility scores fluctuate on random model variations and offer no stable foundation for reporting.
  2. Counted metrics are signal. AI Overview saturation and citation presence provide deterministic, census-style data that holds direction across reporting periods.
  3. Click-level attribution has collapsed. Because referrer data is routinely stripped, marketers must build a multi-layer proxy stack to track real impact.

Read on for a detailed deep dive into the data, the exact numbers, and the reproducible calls you can run against your own domain.

Thesis 1: The AI visibility score your vendor sells you probably samples a random process

Every AI visibility percentage you’ve read in a case study, including ones you may have quoted in a pitch, stays unverifiable until the window, the prompt count, and the variance sit beside it. That includes the numbers you’ve been forwarding to your board.

Here’s the evidence of this from 40 controlled runs on one brand.

Metric (40 runs, same 25 prompts) Mean Standard Deviation Range
Visibility score 32.2 10.3 12 to 54
Citation percentage 46.1 17.9 16 to 88
Mention percentage 18.3 8.6 4 to 33

Source: 40 consecutive ChatGPT probe runs against data-mania.com, April 3 to September 6, 2026.

The coefficient of variation is 32%, and the trend across the full 5 months comes in at r² = 0.00. The score wanders around a mean and stays there.

Two metrics for the same brand over 5 months: the AI visibility score oscillates with no trend, while AI Overview saturation climbs steadily

A real change would have to clear roughly 20 points before anyone could tell it apart from random variation at 95% confidence. Most vendors present a 5-point monthly move as a win.

The mechanism is simple arithmetic. My score averages 2 percentages computed on 25 prompts, so a single prompt that flips moves my composite by 2 points. That 12 to 22 jump came from 4 prompts out of 25.

What to do with this: Ask every vendor for the prompt count behind their score, and treat the answer as the first gate. Then ask any agency that’s pitching you an AI visibility case study for their measurement window and their variance. The request is reasonable, and how they handle it tells you most of what you need to know.

Thesis 2: Your AI visibility score describes your tooling as much as it does your brand

The AI visibility score stays inside the vendor that produced it. Any AI visibility figure in a board deck is a statement about a prompt set, so comparing your score against a competitor’s published score produces nothing usable.

I run 2 tools against data-mania.com. On September 6, one reported 22. The other, probing the same domain in the same week, reported 4.

Same brand, same period, and a headline figure differing by more than 5 times. Both tools report accurately on their own terms, because the category still lacks a shared standard for what “AI visibility” counts as.

3 questions to ask before you believe an AI visibility claim: prompt count, measurement window, and variance

What to do with this: pick one tool and stay on it, and choose on a specific criterion. Pick the one that reads across engines rather than one. A probe that only reads ChatGPT samples a single model’s mood on a single morning, while reading across Google AI Overviews, Google AI Mode, ChatGPT, and Gemini gives you more observations per run. Independent engines drift in different directions on the same day, so the aggregate lands in a tighter distribution. Track these cross-engine signals seamlessly with Semrush Prompt Research & Tracking. That reduces the noise, which is a smaller and more supportable claim than most of this category makes.

Thesis 3: AI search has 2 measurable surfaces, and our industry standardized on the volatile (unreliable) one

This is the assertion that reorganized my entire reporting practice, and it’s the one I’d most like you to take away.

A citation score is a sample. A tool sends prompts, a model generates fresh text, the tool counts what came back. Tomorrow the model generates different text. The variance describes the thing being measured.

AI Overview saturation is a census. It counts what sits on a results page. Pull it twice and get the same answer twice, because everything in the measurement path is already fixed. The sampling, the confidence interval, and the prompt count all drop away.

What this means for you: your key performance indicators should lead with counted metrics and place sampled ones, like your AI visibility score, underneath. Here’s the same brand and roughly the same window, measured both ways.

Month Organic KWs AIO KWs Saturation
2025-07 1,412 906 64.2%
2025-10 1,495 1,026 68.6%
2026-01 2,211 1,704 77.1%
2026-04 2,001 1,755 87.7%
2026-08 1,336 1,252 93.7%

Source: Semrush resource_rank_history, data-mania.com, us database, September 2026.

Across all 14 months: r² = 0.964, slope +2.6 points per month, coefficient of variation 14%. Set that beside the r² of 0.00 and the 32% variation from Thesis 1.

One brand, one period, 2 metrics. One climbs cleanly enough to put in front of a board. The other exhibits noise.

How to spot a metric you can report on: a noisy sampled line versus a clean counted trend across 12 or more readings

Steal this prompt:

“Using the Semrush MCP, run the resource_rank_history report for target yourdomain.com in the us database, with export_columns: date, organic_keywords, serp_ai_overview_keywords, serp_ai_overview_positions. Sort date_desc, limit 24. Then calculate monthly AI Overview saturation as AIO keywords divided by organic keywords, fit a linear trend, and report the slope per month, the r², and the coefficient of variation. Flag whether my organic keyword count is shrinking over the same period.”

One call returns the whole series. Ask for that last check explicitly, because a shrinking footprint can lift saturation through arithmetic alone. Mine fell from 2,551 keywords in February to 1,336 in August, which explains a slice of my climb while leaving most of a 30-point move unaccounted for.

Thesis 4: AI Overview saturation is no longer a differentiator

“We’re heavily exposed to AI Overviews” stopped being a meaningful finding. If your category looks anything like mine, everyone in it is exposed, so that line belongs in your context slide rather than your headline.

Which number belongs in which slide: counted metrics lead the report, sampled metrics support it

Here’s my organic competitor set, sorted by keyword footprint.

Domain Organic KWs Saturation In AIO Presence
growthmethod.com 3,116 94.1% 290 9.9%
productled.com 2,399 93.9% 171 7.6%
gtm8020.com 1,301 93.7% 70 5.7%
data-mania.com 1,252 96.5% 26 2.2%
rightsideup.com 1,197 88.0% 34 3.2%
growth-division.com 1,089 94.3% 36 3.5%
geisheker.com 1,002 97.8% 144 14.7%
movingminds.io 589 87.6% 42 8.1%
rankedcmo.com 133 97.0% 6 4.7%
digital-hunch.com 96 97.9% 3 3.2%

Source: Semrush domain_rank, us database, September 2026. Competitor set from Semrush domain_organic_organic.

Every domain lands between 87.6% and 97.8%. The correlation between footprint size and saturation comes in at −0.13, which is effectively zero, so a 3,116-keyword site and a 96-keyword site are equally covered.

What to do with this: Move saturation into your context slide and out of your headline. It establishes the condition of the category, and the strategic question now sits one column to the right. Benchmark your entire market using the Semrush AI Visibility Toolkit’s Competitor Research to pinpoint exact share gaps.

Thesis 5: Citation presence is the open competitive gap, and it’s countable

Citation presence is where budget and content decisions should point, because it varies enormously between competitors who look identical on every other AI search metric. It also carries a margin of error of zero, which a sampled AI visibility score can never offer, and that makes it the rare AI search number that you can build a quarterly target on. You can evaluate your brand’s exact presence and citation reach with Semrush Brand Performance analytics.

Look at the last column of that table above again. Presence runs from 2.2% to 14.7% across domains with matching saturation, a spread of nearly 7 times. geisheker.com works with a smaller footprint than mine and appears inside AI Overviews 7 times as often.

The distinction lives in 2 separate columns that most teams read as one.

  • serp_ai_overview_keywords counts pages carrying an AI Overview.
  • serp_ai_overview_positions counts pages where you sit inside it.
Your AI Overview data has two columns: saturation (96.5% for data-mania.com) versus presence (2.2%)

My own split: 1,252 organic keywords, 1,208 of them on a page carrying an AI Overview, and 26 where I appear as a source. On 96% of the pages I rank on, Google writes an answer above me, and I’m quoted in roughly 1 in 50.

I sit second from the bottom of my own category on the one metric here that holds steady. I’m publishing that because the alternative is picking the window that flatters me, which is the behavior this whole piece argues against.

Steal this prompt:

“Using the Semrush MCP, run the domain_rank report for target yourdomain.com in the us database, with export_columns: organic_keywords, serp_ai_overview_keywords, serp_ai_overview_positions, serp_people_also_ask_keywords. Then give me AI Overview saturation (AIO keywords ÷ organic keywords), AI Overview presence (AIO positions ÷ AIO keywords), and PAA footprint, all as percentages.”

Both AI Overview columns come back labeled identically as “AI overview.” They arrive in the order you requested them, keywords first and positions second, so parse by position rather than by header.

Steal this prompt to see where you stand in your category:

“Using the Semrush MCP, first run domain_organic_organic for yourdomain.com in the us database to list organic competitors with relevance scores and shared keyword counts. Then run domain_rank for each of the top 8, with export_columns organic_keywords, serp_ai_overview_keywords, serp_ai_overview_positions, and build a table of saturation and presence sorted by keyword count.”

Read the relevance scores before you trust the set. Relevance comes from keyword overlap, so a long-tail footprint can score low against domains that really are your peers. Swap in the ones you compete with, then publish how you chose them.

Steal this prompt to turn the percentage into a worklist:

“Using the Semrush MCP, run resource_organic for yourdomain.com in the us database with export_columns: keyword, position, volume, url, triggered_serp_features, domain_serp_features. Sort by traffic_desc, limit 100. Then show me only the rows where the triggered features include an AI Overview and my domain features sit outside it, sorted by search volume descending.”

That last one names actual pages, which makes it the most useful output in this entire article.

The measurement stack these 5 theses produce

Each layer answers one question well.

The 4-layer AI measurement stack: query characteristics, the market layer, the citation layer, and self-reported attribution

Layer 1. Query characteristics, from Search Console. Classify queries by their inherent characteristics rather than topic and keep an eye on the long-form conversational share. Mine moved from 9.7% to 24.4% in 2 months.

Layer 2. The market layer, from Semrush. Saturation, presence, and People Also Ask footprint for you and your category. Get a complete high-level snapshot using the Semrush AI Visibility Overview, including the counted layer, and the one your reporting should lead with.

Layer 3. The citation layer, read across engines. Mentions, citations, and share of voice inside real AI answers (the inputs behind your AI visibility score), read as a trend line across many runs with the prompt count published beside it.

Layer 4. Self-reported attribution. A how did you hear about us field on your forms, read monthly. Imprecise, self-selected, and better than your referrer data right now.

Steal this: 5 rules for reporting AI search

  1. Publish your prompt count beside every AI visibility score. A score computed on 25 prompts moves 4 points when a single prompt flips.
  2. Report trend lines across many runs. One reading says very little about a metric carrying a 32% coefficient of variation.
  3. Label every proxy as a proxy. Branded search lift and direct traffic ratios are correlation, so say that out loud, in the deck.
  4. Keep click-level attribution out of your promises. The referrer header gets stripped before the click reaches your site, so presence, trend, and competitive context are what this category delivers.
  5. Publish your windows. A result comes with a comparison period and a variance attached. Everything else is an anecdote with a percentage on it.

Your dashboard is what earns you the meeting. Your clarity about what it counts and what it estimates is what earns you the renewal.

Resources and tools

  • Semrush AI Visibility Toolkit to master your entire proxy stack (Visibility Overview, Brand Performance, Competitor Research, Prompt Research & Tracking, and AI Search Site Audit). I use it across every client engagement, and I earn a commission when you subscribe through that link.
  • Semrush AI Search Site Audit (part of the AI Visibility Toolkit) to confirm AI crawlers reach your content. Run it before you commit a quarter of production to pages that models skip.
  • Google Search Console for Layer 1 query-characteristics classification.
  • The Semrush MCP server connected to Claude, which is how I run every call above conversationally rather than through exports.

Semrush sponsored this article. The analysis, the data, and the conclusions are mine, and every figure here is reproducible against your own domain using the prompts above. This post also contains affiliate links.

Best regards,

Lillian

P.S. The morning my score read 22, I almost screenshotted it. Up 83% in 4 days, and it would have made a beautiful LinkedIn post. I had the crop selected.

A habit from my engineering years stopped me, where you check the instrument before you celebrate the reading. Twenty minutes later I had a spreadsheet showing me that my beautiful 83% was a Tuesday.

If you’d like this kind of measurement discipline engineered into your own growth system, I write about it every week in The Convergence, my newsletter portfolio with over 100,000 subscribers. Or book a 30-minute conversation and we’ll map your measurement stack together.

Share Now:

TURN YOUR GROWTH GAPS INTO PROFIT CENTERS

From roadblocks to revenue: it all starts here. Get your free Growth Engine Audit & Gap Map™ now to uncover the tangible growth opportunities that are hiding in plain sight.

IF YOU’RE READY TO REACH YOUR NEXT LEVEL OF GROWTH