// GUIDE · 2026-08-13

How AI search visibility metrics are actually calculated (2026): the formulas, raw inputs, and benchmarks behind AI share of voice, citation rate, and prompt coverage

Every team that has watched an AI Overview eat its organic clicks now agrees it needs "AI search visibility metrics" — share of voice, citation rate, prompt coverage, prominence, sentiment. Far fewer can say how any of them is calculated, which is a problem, because two vendors can hand you a number called "AI share of voice" that were computed from different denominators and are not comparable. This guide is the mechanics, not the pitch: it starts from the one raw input every metric derives from — a fixed prompt set run repeatedly across each engine and logged answer by answer — and then works through the arithmetic of each headline metric one at a time. What the numerator and denominator actually are for prompt coverage versus true share of voice (they are different numbers that share a name), how citation rate and citation share differ and which one answers "who won this query," how to score prominence when there is no ranked list to read a position off, and why sentiment and AI-referral traffic are the two you can compute but should not chase. It is honest about the parts nobody can hand you cleanly: there is no industry-standard benchmark for most of these yet, the measurement is non-deterministic so any single reading is a sample rather than a fact, and per-engine numbers must never be blended into one score. It closes on the one term in every one of these formulas that you actually control — the published content the engines can cite — and how to keep enough of it in supply to move the number.

Last verified · 2026-08-13 · by Moe Ameen

Everyone agrees they need the metrics; almost nobody defines them

The case for tracking AI search visibility is settled. When roughly half of Google searches return an AI Overview and organic click-through-rate on those queries falls by roughly 61% in Seer Interactive's measurement, a page can hold its ranking and lose its traffic in the same month, so the old scorecard stops tracking demand. That much has been argued to death, including in the companion pieces on running AI visibility as a growth channel and what to track versus ignore. This guide assumes you are past the why and stuck on the how.

Because here is the gap. Ask three vendors for your "AI share of voice" and you can get three different numbers, all honestly reported, all computed from different denominators, none comparable. "Citation rate" and "citation share" get used interchangeably and mean different things. "Position" is quoted as if it were a keyword rank when there is no ranked list to read it off. The metrics have outrun their definitions, which is dangerous on a number you are about to put in a board deck or steer a content budget by. So this is the arithmetic, one metric at a time, starting from the single input they all come out of.

The one input everything derives from: the answer run matrix

None of these metrics is measured directly the way a server logs a pageview. They are all computed after the fact from one dataset, and understanding that dataset is most of the battle. You build it by running a fixed prompt set across each engine, repeatedly, and logging every answer in a structured way. That log — call it the answer run matrix — is the raw material; every headline metric is just a different slice of it.

Concretely, a run is one prompt sent to one engine one time. For each run you record: the prompt, the engine (ChatGPT, Gemini, Perplexity, Google's AI Mode, Copilot, Claude — kept separate, never merged), the timestamp, the full answer text, whether your brand was mentioned, which competitor brands were mentioned, the order or placement of each mention, and the list of source URLs the answer cited. Because answers are non-deterministic, you run each prompt several times per engine and treat the set as a sample, not a single reading. The two upstream decisions that govern everything are the prompt set itself — 20 to 30 questions mined from real sales calls and support tickets, not invented from how you describe yourself — and running enough repetitions that a single lucky or unlucky answer does not swing the number. Every metric below is arithmetic on this matrix, so if the prompt set is wrong or the sample is too small, no formula downstream can rescue it.

Prompt coverage (appearance rate): how often you show up at all

The most basic metric is also the most misreported, because it is the one people mean when they loosely say "share of voice." Prompt coverage — also called appearance rate or visibility rate — is the fraction of runs in which your brand is mentioned at all. The formula: coverage = (runs that mention your brand) ÷ (total runs), computed per engine. If you run 30 prompts five times each on Gemini, that is 150 runs; if your brand appears in 45 of them, your Gemini coverage is 30%.

Two disciplines make this a metric rather than a vibe. First, keep it per engine and never blend — the same prompt returns different brands on ChatGPT and Gemini, and a blended figure hides exactly the surface you are losing. Second, decide up front whether "mentioned" means a bare name-drop or a substantive mention, and apply that rule consistently, because a tool that counts any passing reference will report a rosier coverage than one that requires a real recommendation. Coverage answers one narrow question well: when someone asks the engine about this category, how often does your name come up? It says nothing about whether you came up first or whether a competitor came up more, which is what the next two metrics add.

AI share of voice — and the two different numbers that share the name

This is the term most likely to cause a bad decision, because it has two legitimate definitions that produce different numbers, and vendors rarely tell you which one you are looking at. Classic marketing share of voice, from the PR and social-listening world, is always competitive: your mentions as a fraction of all mentions in the category. Applied to AI answers, that gives you the honest, competitive reading. But many AI-visibility tools relabel simple appearance rate as "share of voice," which is a presence number, not a competitive one. Same words, different math.

The presence version (really coverage)

This is prompt coverage from the section above, sometimes dressed up as "share of voice": (prompts where you appear) ÷ (total prompts), per engine. It is useful and easy, but it is not share of anything — it does not look at competitors at all. A brand can have 60% coverage and still be the least-mentioned name every time it appears. If a tool's "share of voice" number goes up when you publish more and never moves when a competitor surges, it is measuring presence, not share.

The true share-of-voice version (competitive)

The competitive version divides your brand's mentions by the total of all brand mentions across the tracked runs: SOV = (your mentions) ÷ (your mentions + every competitor's mentions). Run 150 prompts, count every branded mention in every answer, and your share is your slice of that total. This is the number that behaves the way share of voice is supposed to — it falls when a rival gets mentioned more even if your own coverage held steady, which is precisely the competitive signal the presence version cannot give you. When you report or consume an "AI share of voice" figure, the one question that makes it meaningful is: what is the denominator — total prompts, or total brand mentions? Do not compare two tools' numbers until you know that both used the same one.

Citation rate and citation share: two different competitive questions

Coverage and share of voice count whether your name appears in the prose. Citations count something more concrete and more valuable: whether the engine actually pulled from your website as a source, shown as a linked reference under or within the answer. This is the closest AI search gets to "who won this query," and like share of voice it splits into two metrics that are routinely confused.

Citation rate — how reliably you get sourced

Citation rate measures you against the answers: citation rate = (answers that cite your domain) ÷ (total answers), per engine. Of every answer the engine produced for your prompt set, how many reached for your site as a source? It is the retrieval-side companion to coverage — coverage asks whether you are named in the text, citation rate asks whether you are footnoted as a source, and the two can diverge sharply. A brand can be described in an answer without being cited, or cited without being named prominently.

Citation share — how much of the sourced material is yours

Citation share measures you against everyone cited: citation share = (your cited URLs) ÷ (all cited URLs across those answers). If a set of answers cited 200 URLs and 24 of them were yours, your citation share is 12%. This is the directly competitive one — it tells you how much of the material the model built its answers from came from you versus rivals versus neutral third parties, and it is the most persuasive single slide because it maps to a concrete editorial target: a topic cluster where competitors own the citations is a gap you can name and staff. State the caveat every time, though: a citation is not a recommendation. Being one of six sources under an answer that ends up recommending a competitor is a citation, not a win, so read citation share next to prominence rather than alone. Which content formats actually earn the citation is worked out in why specific, niche content gets cited more, and the retrieval mechanics behind who gets pulled are in how Perplexity selects sources.

Prominence: scoring position when there is no ranked list

Classic SEO had position because there were ten ordered blue links. A generative answer is a paragraph, so there is no clean coordinate to read a rank off — but placement still varies, and it matters, so prominence is the metric that tries to capture it without pretending there is a leaderboard. The workable approach is to score each mention on a small ordinal scale: named first or as the primary recommendation, named among several, or mentioned only in passing after competitors. Average those scores across the runs where you appear, and track the average as a trend line, not a precise number.

Two honest limits keep this from being over-engineered. First, it is only computed over runs where you were mentioned at all — prominence and coverage answer different questions and must not be collapsed. Second, there is no universal scale, so whatever ordinal buckets you choose, keep them fixed month to month or the trend becomes meaningless. A brand that moves from "named last" to "named first" on its priority prompts is winning even when raw coverage barely moved, and prominence is the only metric that catches that shift.

Sentiment and AI-referral traffic: computable, but not to be chased

Two more numbers are worth computing and worth being disciplined about, because both are easy to over-weight. Sentiment scores how the engine characterizes you when it mentions you: tag each mention positive, neutral, or negative, and — separately and more urgently — flag any statement that is simply factually wrong, because models confidently repeat outdated pricing or a discontinued product. The sentiment split is a thermometer you watch, not a dial you turn; the factual-error flag, though, is often the single most valuable thing a visibility audit surfaces, because fixing the source content that feeds the error is concrete work.

AI-referral traffic is the weakest metric of the set, and it is worth saying why so you do not lean on it. Many AI answers include no clickable link at all, platforms change how they pass referrers without notice, and the number swings for reasons you did not cause, so a low figure understates your real influence and a moving figure rarely means what it looks like. Because so much AI influence never produces a traceable click, the more reliable off-platform proxy is a durable rise in branded search and direct navigation that tracks your publishing — imperfect, since brand demand has many causes, but real when it trends with your topic coverage. The cleanest way to tie any of it to revenue is the low-tech one every attribution guide lands on: a "how did you hear about us?" field on your forms, because standard analytics mislabels most AI-influenced leads as direct or organic. Google's own emerging AI-surface impression data helps on its side specifically, covered in reading Search Console's AI impressions metric.

Reading the numbers honestly: no benchmark, and every reading is a sample

Two properties of these metrics change how you are allowed to read them, and skipping them is how teams draw wrong conclusions from correct math. The first: there is no reliable cross-industry benchmark for what a "good" score is, and there will not be one soon. The metrics are barely a year or two old, the engines and their sourcing shift monthly, and every number depends on your specific prompt set, your category's competitiveness, and which engines you weight. A vendor's suggested threshold is a marketing artifact, not a standard. The only benchmark that means anything is your own baseline measured on your own fixed prompt set — where you started, and whether you moved.

The second: the measurement is non-deterministic. Ask the same engine the same question twice and it can name different brands and cite different sources, so any single run is a draw from a distribution, not a fixed value like a keyword rank. That is exactly why the answer run matrix is built on repeated runs on a schedule, why per-engine figures are never averaged into one blended score, and why you read a trend over months instead of reacting to one answer. Present that up front to anyone consuming the numbers — a stakeholder who understands the sampling reads the trend line correctly instead of panicking at a single bad screenshot. This is the same discipline the client-reporting playbook and the content-marketing scorecard both build their monthly rhythm around, and it is the frame the whole field of generative engine optimization now works inside.

The one term in every formula you actually control

Read the formulas back to back and a pattern jumps out. Coverage, share of voice, citation rate, citation share, prominence — every one of them has your published content sitting in the numerator, and it is the only term on the page you control. You cannot force an engine to sample your prompt, cannot set the denominator of competitor mentions, cannot make a model cite you. You can change exactly one thing: whether a genuinely useful, specific answer to that prompt exists, on a surface the engine crawls, in enough supply and refreshed often enough to survive the sampling noise. Every measurement tool stops precisely here — it computes the score and hands you a backlog of prompts you are losing, and it writes and publishes nothing. That is the production gap the number correctly exposes, not a measurement failure.

Kompozy is built for that numerator. It is an AI content generation and multi-platform publishing engine — pair it with whichever tool owns your scoreboard; Kompozy is how you put points on it. You take one losing prompt the audit surfaced and a real answer to it — a subject-matter voice memo, a support doc, a page that ranks but no longer earns the click — and the engine produces across several output buckets at once: a full blog article that answers the prompt directly, plus text posts, images and carousels, and short-form or avatar video, each shaped to a different surface an engine samples. Breadth across surfaces is what raises coverage; specificity is what the citation formulas reward, since detailed niche pages get pulled ahead of generic ones.

The parts that make the metrics move at scale are the parts a lean team cannot hold by hand. Every generation descends from one Persona Brief — voice, point of view, banned phrases — so producing enough answers to beat the sampling noise stays specific and citable per brand instead of collapsing into the median filler models skim past. Autopilot then schedules the finished set across eight social platforms plus blog and email behind a per-post human review gate, so the answers land on the surfaces the engines actually read rather than sitting on one page. The honest boundary is the same one to state to any stakeholder: Kompozy cannot manufacture expertise your business does not have, cannot decide which prompts are worth owning, and cannot make any engine cite anyone — citation is the model's call. What it removes is the production ceiling that otherwise caps you at a post a week when the audit says a dozen gaps need answering. Measure with the tool; supply the numerator with the engine; re-run the same prompt set next month and read the movement.

Frequently asked questions

What are the core AI search visibility metrics, and what does each measure?

Five recur across every tool. Prompt coverage (appearance rate): the share of your tracked prompts where you are mentioned at all, per engine. AI share of voice: your brand's mentions as a fraction of all brand mentions across the same runs — a competitive number, not just presence. Citation rate: the share of answers that link to your domain as a source. Citation share: your cited URLs as a fraction of all cited URLs, the closest thing to 'who won this query.' Prominence: where you land when you are named — first, buried, or in passing. Sentiment sits alongside as a qualitative sixth.

How is AI share of voice actually calculated?

It depends which definition a tool uses, and they are not the same number. The presence version — often labelled visibility or appearance rate — is (prompts where you are mentioned) divided by (total prompts run), per engine. The true share-of-voice version is (your brand mentions) divided by (all brand mentions across every tracked run, yours and competitors'). The first tells you how often you show up; the second tells you how much of the category conversation is yours versus rivals'. A vendor reporting 40% could mean either, so always ask for the denominator before comparing two tools' figures.

What is the difference between citation rate and citation share?

Citation rate is about you against the possible answers: how many of the answers in your run actually linked to your domain, i.e. (answers citing your domain) divided by (total answers). Citation share is about you against everyone cited: your cited URLs divided by all cited URLs in those answers, which measures competitive footing on the sources the model reached for. Rate tells you how reliably you get pulled in; share tells you how much of the sourced material was yours. Both matter, and neither is a recommendation — being one of six links under an answer that recommends a rival still counts as a citation.

Is there a benchmark for a "good" AI visibility score?

Not a reliable cross-industry one, and treating a vendor's suggested threshold as a standard is a mistake. These metrics are only a year or two old, the engines and their sourcing change monthly, and the numbers depend entirely on your prompt set, your category's competitiveness, and which engines you weight. The only benchmark that means anything is your own baseline: run a fixed prompt set, record where you start, and measure movement against last month. Comparing your appearance rate to a number from a different prompt set or a different engine mix compares two things that were never the same measurement.

Why can I not treat one AI visibility reading as a fixed number?

Because generative answers are non-deterministic. Ask the same engine the same prompt twice and it can name different brands and cite different sources, so any single run is a sample from a distribution, not a fixed value like a keyword rank. That is why the metrics are built on a prompt set run repeatedly on a schedule rather than one snapshot, why per-engine numbers are never averaged into one blended score, and why you read the trend over months instead of reacting to a single good or bad answer. The sampling noise is a property of the medium, not a flaw in the tool.

How does Kompozy help move these numbers rather than just measure them?

Every one of these formulas has exactly one term you control: the published content the engine can mention or cite. A tool computes the score; it cannot raise the numerator. Kompozy is an AI content generation and multi-platform publishing engine that manufactures that numerator — from one source it produces a blog article answering a losing prompt directly, plus the text posts, carousels, images, and short-form or avatar video that carry the answer onto the surfaces engines sample, all held to one Persona Brief and scheduled across eight social platforms plus blog and email behind a per-post review gate. It supplies the volume and specificity the citation formula rewards; you supply the expertise and the final yes.

The direct answer

AI search visibility metrics all derive from one raw input: a fixed prompt set run repeatedly across each engine, logged answer by answer for whether your brand is mentioned, where it lands, and which domains are cited. From that log, prompt coverage is appearances divided by total runs; AI share of voice is your mentions divided by all brand mentions; citation rate is answers citing your domain divided by total answers; citation share is your cited URLs divided by all cited URLs. Compute each per engine, and read every figure as a sample, not a fixed rank.

Get started → · ← All guides · Compare Kompozy vs other tools