Industry Insights · · 8 min read

How to Measure AI Visibility: Mentions, Citations and Coverage

An illustrative prompt-by-provider table separates brand mentions, cited links and unavailable responses.

AI visibility measurement records what a fixed set of answer systems returned for a defined set of questions, on the day you asked them. It is a sample, not a census of everything every customer sees. The value of the exercise comes from what you keep alongside the headline number: the exact question, the exact answer, and the source evidence, so that a second reviewer can check your working rather than take your word for it.

This guide sets out a practical way to build that record, calculate a rate that means something, and avoid the most common ways these reports mislead the people reading them.

Define the sample before you run it

Start by choosing questions from real customer decisions, not from whatever comes to mind on the day. Separate three types of question:

  • Category questions that do not name any brand, for example "which inventory tools suit a small retailer?"
  • Questions naming the client's own brand, for example "what does Example Inventory cost?"
  • Questions naming a named competitor

These are different tests. A category question checks whether the brand appears when the question does not name it. A branded question mostly checks whether the assistant has basic facts about a business that already told the assistant its own name. Mixing the two in one blended "mention rate" can make the overall number look stronger without any real gain in category visibility, which answers a different question about discovery.

Consider a clearly fictional inventory-software company that appears in answers explicitly naming it, but is absent from category recommendations. Combining those prompt classes can conceal the category gap. Report them separately, and avoid describing a branded lookup as an unsolicited recommendation.

Once you have chosen your prompt classes, keep the prompt text, market, language and provider set stable whenever you intend to make a time comparison. If you change the wording of a prompt, add a market, or drop a provider, treat that as a break in the series and say so in the report rather than plotting it as a smooth line. Writegarden's current implementation includes ChatGPT, Perplexity, Gemini, Copilot, Claude and Google AI Overviews, but you should inspect the actual response coverage for each run rather than assume every provider returned a usable answer on every date.

A sample separates delivered answers from missing responses before measuring mentions.

Build an observation sheet you can defend

Every run should produce a row per prompt per provider, with enough detail that someone who was not in the room can understand the result. The fields below are the minimum that makes a report defensible rather than merely plausible.

FieldPurpose
Prompt ID and exact textLets you reproduce the question later and catch silent wording changes between runs
Prompt classKeeps category, branded and competitor-branded questions in separate buckets
Provider, market, language, dateRecords the exact conditions the observation was made under
Response statusDistinguishes delivered, pending, failed and unavailable results, so a gap is not silently read as an answer
Original answerKeeps the underlying evidence, not just a label someone attached to it
Brand mentionedRecords presence as a yes/no, kept separate from how many times the name appears
Cited URLsLists any linked sources, flagging which ones belong to the client
Reviewer noteExplains ambiguous brand names, outdated claims in the answer, or any classification you had to judge by eye

A failed collection or an absent AI Overview is not proof that a brand lost visibility. It is a missing observation. Record the status honestly and go and check whether the collection itself failed before treating a blank row as a lost mention.

Calculate a mention rate with a visible denominator

For any defined group of prompts, the mention rate is: delivered responses mentioning the brand, divided by delivered responses in that same group, multiplied by 100. Always report both counts next to the percentage, not the percentage on its own.

Here is a worked example, illustrative only and not a record of any client's actual performance. A scan requests 30 responses, receives 24 usable answers, and finds the brand mentioned in 6 of them. The mention rate is 6 divided by 24, which is 25%. Coverage, meaning how many of the requested responses actually came back, is 24 divided by 30, which is 80%. The 6 responses that never returned an answer are unknown, not six observed absences of the brand. Treating them as absences would understate the rate and overstate confidence in a number built on a fifth of the sample being missing.

In the Writegarden implementation reviewed on 22 September 2026, delivered responses are aggregated and separate rates are calculated by provider and by prompt class. The main category view uses category prompts when that metric is available; older snapshots may use the overall rate instead, so a chart spanning both periods can quietly switch what it is measuring halfway along. Some empty provider buckets display a numeric zero rather than a blank, which looks identical to a genuine zero mentions result. Always check the response count and status behind a zero before reporting it as a real finding, and compare only matching metric definitions when you line up two dates.

Two answer cards distinguish a brand mention from a link citing a source page.

Keep citations, sentiment and causation separate

A response can mention a business by name without linking to any of its pages. It can equally cite a client's page without ever repeating the company name in the visible text. Record both of these as separate observations rather than folding them into one score. Count the number of responses containing a client citation separately from the total number of links across all responses: three links inside one answer are not three independent citing answers, and reporting them as such inflates the picture.

Sentiment labels are a reading aid for a human reviewer, not a verified judgement. The current classifier works from keyword heuristics, which can miss negation, context, sarcasm and differences between languages. Read the original answer before putting a sentiment label in front of a client. The marketing name Harvest Index should not be treated as a single verified citation-and-sentiment weighting; the current implementation reports citation counts and sentiment separately, and a combined score is a description, not an audited metric.

When a mention rate moves between two runs, resist the temptation to attach a cause to it in the same paragraph. Preserve the baseline period, any content changes made in between, the prompt sample and the provider coverage, then repeat the observation and read the underlying answers before drawing a conclusion. A mention-rate increase on its own does not establish that a new article, a schema change or an outreach push caused the increase; several things usually changed at once, and the assistants themselves change their retrieval behaviour independently of anything a marketing team did.

Measure website referrals, Search Console performance and app conversions as their own separate exercises. They answer different questions from a sampled answer test and should not be merged into one number to imply that AI visibility explains a change in traffic. Google's AI feature guidance remains the primary reference for its own Search surfaces; other providers have their own retrieval and citation behaviour and should not be assumed to follow the same rules.

Common mistakes worth checking for before you send a report

  • Blending category, branded and competitor-branded prompts into a single rate, which hides whether the brand appears in category answers or only when already named
  • Reporting a percentage without the underlying counts, so a reader cannot distinguish a small sample from a larger one
  • Treating a failed or unavailable response as a confirmed absence rather than a missing observation
  • Comparing two dates that use different metric definitions, for example an overall rate against a category-only rate
  • Counting several links inside one answer as several separate citing responses
  • Quoting a sentiment label without reading the original answer it was generated from
  • Claiming a specific content change caused a mention-rate shift without checking what else changed over the same period
  • Changing prompt wording, market or provider set between runs without flagging the break in the series

Decision criteria: when to trust the number, and when to investigate first

Before acting on a mention rate, check that coverage is reasonably high for the group it describes; a rate built on a handful of delivered responses out of a much larger requested sample deserves a caveat, not a confident headline. Check that the prompt class matches the claim being made, so a branded-question result is never presented as evidence of category visibility. Check that the date range compares like with like, using the same prompt set, market, language and metric definition throughout. If any of these checks fail, treat the number as a starting point for further reading of the original answers, not as a finished finding.

Use this method alongside the AEO brief and the technical page checklist. A report earns trust when the evidence and its limitations stay visible on the page, even when the sample is small, the coverage is incomplete, or the result is less flattering than hoped.

About the author

WriteGarden

WriteGarden is an AI-native SEO and AEO content operations platform, a product of Digital JATO.

Plant your first cluster.

€1 trial · Credited to your first invoice · Cancel anytime

Try for €1 →