Growli
Blog

AI Share of Voice: How to Measure and Benchmark It

Zakaria Reziki

By Zakaria Reziki

CEO โ€” Growli ยท August 13, 2026 ยท 11 min read

Drafted with AI assistance under editorial standards set by Zakaria Reziki, then published after automated sourcing and quality checks.

AI share of voice is the percentage of tracked prompt runs on a given AI assistant in which your brand is mentioned โ€” brand mentions รท total prompt runs, calculated per assistant rather than blended into one figure. It is the closest equivalent to organic share of search in the answer-engine era, and like share of search it is only meaningful when the question set, the sample size and the counting rule are held constant.

The metric matters because the click is no longer guaranteed. Pew Research Center found that Google users who encountered an AI summary clicked a traditional search result on 8% of visits, compared with 15% of visits without one (Pew Research Center). When the answer replaces the result page, being named inside the answer is the impression.

This piece defines the formula, explains why per-assistant benchmarks legitimately diverge, gives a table template you can copy into a sheet today, and shows how to convert each type of gap into a specific action rather than a vague content brief.

The metric, stated precisely

Most teams conflate two different things under the same label. Separate them and the diagnosis gets much easier.

Mention rate answers a binary question: do you show up at all? Share of answer answers a competitive one: of all the brands the assistant named, how much of the airtime was yours? A B2B tool can have a 90% mention rate and still lose, because it is named ninth in a list of ten and the first three get the buyer's attention.

Both are calculated per assistant and per prompt cluster. A cluster is a group of buying-intent questions that belong to the same decision โ€” pricing comparison, alternatives to a named competitor, best-tool-for-X, integration questions. Averaging across clusters produces a number that moves for reasons you cannot trace.

  • Mention rate (%) = runs in which your brand appears รท total prompt runs, for one assistant and one cluster, ร— 100.
  • Share of AI answer (%) = your brand mentions รท all brand mentions across the same answers, ร— 100. This is the true share metric because it is zero-sum against a defined competitor set.
  • Weighted share (optional) = the same ratio with each mention weighted by position, so being first in a list counts more than being a footnote. Use a simple weight (first = 1.0, mid = 0.6, last = 0.3) and document it.
  • Recommendation rate = runs where you are named as a recommendation, not just referenced as an example. Worth splitting out for high-intent prompts.

Why it matters

Clicks on results, with and without an AI summary

Source: Pew Research Center

The three inputs that decide whether your number means anything

An AI visibility metric is only as stable as its instrumentation. Three decisions do most of the damage when they are made loosely.

First, the prompt set. Fix it in writing โ€” the exact phrasings, the market, the language, the persona framing โ€” and version it. Every time you add prompts, your trend line breaks, so add them as a new cluster rather than editing the old one. Aim for enough prompts to cover the decision, not enough to look thorough: twenty to forty well-chosen buying questions per market beat several hundred generic ones.

Second, the run count. Assistant outputs are not deterministic; the same prompt asked twice can name different brands, especially where live retrieval is involved. That means a single run is an anecdote. Run each prompt multiple times per assistant per period โ€” five is a workable floor, more if your category has a long tail of competitors โ€” and treat small gaps on small samples as noise rather than movement.

Third, the mention rule. Decide in advance whether a bare URL counts, whether a citation in a source list counts as a mention, how you handle misspellings and parent-brand names, and whether a negative mention counts. Write the rule down; inconsistency here creates most of the disputes about whether visibility improved.

Why per-assistant benchmarks diverge โ€” and why that is correct

A brand at 55% share of answer in Perplexity and 12% in ChatGPT for identical prompts is not a measurement error. The systems differ at the layer that decides which brands are eligible to be named.

Retrieval versus parametric knowledge is the biggest split. Assistants that lean on live web search surface whatever the current index and their ranking layer favour; answers that come mostly from model weights reflect how prominently a brand appeared in training data. OpenAI describes ChatGPT search as combining the model with information from the web and third-party providers (OpenAI), Anthropic ships web search as a tool Claude decides when to invoke (Anthropic), Perplexity is built around retrieval and citation by default (Perplexity), and Google documents its AI features as a surface layered on Search with its own linking behaviour (Google Search Central). Different eligibility, different brand lists.

Source preference is the second driver. One assistant may lean on editorial roundups and review platforms, another on documentation, forums and primary sites. If your category's third-party coverage is concentrated in a handful of listicles, the assistants that trust those pages will over-represent whoever ranks well on them.

Freshness and grounding rules are the third. A newly launched competitor can appear immediately in retrieval-heavy assistants and be invisible for months in weight-heavy ones. Region, language and whether a tool decides it needs to search at all add further legitimate variance.

The practical consequence: never report a blended cross-assistant average as your headline. Report a matrix. The spread between assistants is itself the insight, because it tells you whether your problem is source coverage, ranking on other people's pages, or plain absence from the corpus. That per-assistant breakdown is the core of what we build at Growli, and it is the reason a single global score is the least useful number in the report.

Structural, not noise

The three drivers behind diverging numbers

  • Retrieval vs parametric knowledge

    Live-search assistants surface what the index favours; weight-heavy ones reflect training data.

  • Source preference

    Roundups and review platforms for one assistant; documentation, forums and primary sites for another.

  • Freshness and grounding

    A new competitor appears immediately in retrieval-heavy assistants and months later in weight-heavy ones.

A benchmark table template you can copy

Build one row per assistant per prompt cluster per period. Columns, in order:

Read it two ways. Across a row, you see whether a cluster is healthy on one assistant and broken on another. Down a column, you see whether a competitor is winning everywhere (a corpus and authority problem) or on one surface only (a source problem you can target).

Add a second sheet holding the competitor set โ€” five to eight named rivals, fixed for the period โ€” plus the exact prompt list and the mention rule. Without those, next quarter's number is not comparable to this one.

  • Period โ€” week or month, with the run dates.
  • Assistant โ€” ChatGPT, Gemini, Claude, Perplexity, Copilot, Grok. One row each; no blending.
  • Prompt cluster โ€” e.g. best-tool-for-X, alternatives-to-competitor, pricing.
  • Runs โ€” total prompt runs behind the row. Anything under five per prompt should be flagged as low confidence.
  • Your mentions โ€” runs containing your brand under the documented rule.
  • Mention rate % โ€” your mentions รท runs.
  • Total brand mentions โ€” all mentions of you and the tracked competitor set in the same answers.
  • Share of answer % โ€” your mentions รท total brand mentions.
  • Average position โ€” mean rank when named in a list.
  • Top competitor + their share % โ€” who is actually taking the slot.
  • Top cited domains โ€” the three sources the assistant leaned on most; this is your action list.
  • Delta vs previous period โ€” in percentage points, not percent change, so small bases do not exaggerate movement.

Turning SoV gaps into actions

A share-of-voice report that ends in a percentage is a dashboard. The value is in the mapping from gap type to intervention, and there are only a handful of gap types worth naming.

Absent on every assistant for a cluster means you are not in the source material the answers are built from. The work is coverage: get named, described and compared on the pages assistants already cite for that cluster โ€” review platforms, category roundups, community threads, industry publications โ€” and make sure your own pages state plainly what you do, for whom and at what price.

Present in retrieval-heavy assistants, absent in weight-heavy ones points to recency. You exist on the live web but not in the deeper corpus. Expect a lag, and in the meantime concentrate on the surfaces where retrieval decides eligibility.

Named but ranked last is a comparative-evidence problem. Assistants order lists using the framing they find in sources. If competitors have explicit comparison pages, pricing tables and use-case statements and you do not, you get slotted as the also-ran. Publish the specifics โ€” who you are not for is as useful as who you are for.

Named with wrong attributes is the most damaging and the easiest to miss if you only count mentions. Log the claims made about you, not just the fact of the mention: wrong pricing model, a missing integration, an outdated positioning. Correct it at the source that is being cited.

One competitor dominant on one assistant usually traces to a single page. Find it in the cited-domains column, then decide whether the play is inclusion in that page, a stronger alternative on a comparable domain, or a factual correction request.

Prioritise by commercial weight, not by gap size. A ten-point gap on high-intent alternatives-to prompts is worth more than a forty-point gap on definitional questions. If you want the wider framing of how mention rate, sentiment and citation sources fit together, see our primer on AI visibility.

Gap to action

Read the gap type, then pick the fix

  1. Absent everywhere

    You're missing from the source material โ€” earn coverage on the pages assistants cite.

  2. Named but ranked last

    Weak comparative evidence โ€” publish explicit comparisons, pricing and use-case pages.

  3. Wrong attributes

    Log the claims, then correct them at the cited source.

  4. One competitor dominant

    Trace it to the single page doing the work, then target inclusion or an alternative.

Cadence, confidence and what counts as real movement

Monthly measurement suits most categories; weekly is justified during a launch, a rebrand or an active correction campaign. Whatever you pick, run the full prompt set in a tight window so model updates do not straddle your sample.

Set a movement threshold before you report. With five runs per prompt, a shift of a few percentage points on a single cluster is well inside sampling variance โ€” call movement only when it holds for two consecutive periods or appears across multiple clusters on the same assistant. Annotate every report with the events that could explain a jump: a model update, a new competitor launch, a large third-party article going live, a change to your own site.

One last discipline: keep the raw answers, not just the tallies. Six months from now the tallies will tell you that share fell and the stored answers will tell you why โ€” which is the difference between an AI visibility metric you can defend and a number that starts an argument.

See what AI says about your business

Growli measures your share of AI answers across ChatGPT, Gemini, Claude and Perplexity โ€” and turns every gap into prioritized actions.

Get Started

FAQ

There is no universal benchmark, because the number depends entirely on your prompt set and competitor set. The useful baseline is relative: your share of answer versus your closest three competitors on the same prompts and the same assistant, and your own trend across periods. A category leader in a five-brand market might sit above 30% share of answer, while the same figure in a crowded market with thirty credible tools would be exceptional.

Share of search measures the proportion of category search demand that names your brand, based on query volume. AI share of voice measures the proportion of AI-generated answers that name your brand, based on prompt runs you control. Share of search reflects existing demand; AI share of voice reflects which brands the assistant considers eligible to recommend, so it can move even when demand is flat.

Because assistant outputs are not deterministic, one run per prompt is an anecdote. Five runs per prompt per assistant per period is a workable floor for a stable mention rate, and more is better in categories with many near-equivalent competitors. Always record the run count alongside the percentage so readers can judge confidence.

Report a per-assistant matrix with an optional weighted roll-up beneath it, never a blended number alone. Assistants differ in retrieval architecture, source preference and freshness, so a single average hides the pattern that tells you what to fix. If leadership needs one number, weight each assistant by its share of your actual audience and show the components underneath.

That is your decision, but it has to be made once and documented. Many teams count a named brand in the answer body as a mention and track source citations as a separate column, because the two mean different things: the body mention is the recommendation, the citation is the mechanism. Mixing them inflates your score and blurs the action.

Yes, for a small prompt set. Fix twenty prompts, run each five times on two or three assistants, log the brands named and the sources cited, and compute the two ratios in a sheet. It becomes impractical once you need multiple assistants, markets and languages repeated on a schedule with the raw answers stored, which is the point at which teams move to tooling โ€” ours or anyone's.

Weekly newsletter

Weekly AI search growth tips

Join business owners getting practical tips for being found on Google and recommended by AI.