Back to Insights

AI Search Visibility

AI visibility metrics that matter (and the vanity metrics to ignore)

5 min read

An analogue control panel dense with dials and gauges.

The short answer

Ask three questions of any number before it earns a place in an AI visibility report. Is it measured on a fixed set of questions your buyers actually ask? Is it comparable month over month, because neither the questions nor the method changed between readings? And does it connect to something commercial: being cited while a buying decision is researched, or traffic and enquiries arriving from AI surfaces?

Five metrics pass that test: citation rate on a fixed prompt set, share of citations against competitors on those prompts, accuracy of what the engines say, referral traffic and assisted conversions from AI surfaces, and the month-over-month trend on an unchanged prompt set. Most of what gets reported as “AI visibility” fails it: raw mention counts, sentiment scores, one good screenshot, and blended scores with no visible method. The rest of this guide defines both lists and the judgement behind them.

Metrics that matter, and their vanity counterparts

Worth trackingVanity counterpartWhy the vanity version misleads
Citation rate on a fixed prompt set: the share of defined buyer questions on which your domain is citedRaw mention volume with no prompt controlA big number that moves when the prompt mix changes, not when your visibility does
Share of citations versus competitors on those same promptsMentions on prompts nobody commercially asksYou can “win” questions that never precede a purchase
Accuracy of what the engines say about youSentiment scores without citationsPositive but wrong is still wrong, and without the cited sources you cannot verify or fix it
Referral traffic and assisted conversions from AI surfacesOne-off screenshots of a good answerA screenshot is one run of one engine on one day; the next run can answer differently
Month-over-month trend on an unchanged prompt setA composite “AI visibility score” with undisclosed methodologyA number you cannot decompose cannot tell you what to fix, and can move on a definition change

Four of the five deserve a sharper definition.

Citation rate is the core selection metric. Write down the questions a buyer would put to an assistant before shortlisting in your category — real commercial questions, not brand searches — and measure, across repeated runs, the share of those on which your domain appears among the cited sources. It captures the thing that actually matters: when the defined question is asked, are you part of the evidence the answer is built from, or not.

Share of citations is the competitive layer on the same data. Count every citation the engines return across the prompt set, group them by domain, and express yours as a share of the total. It is the nearest thing AI search has to a rank, and it moves for reasons you can inspect: a competitor publishing, a source dropping out, your own pages starting to be selected.

Accuracy is the risk metric. Engines answer with confidence whether or not they are right, and a wrong claim about your pricing, positioning or ownership travels further than any correction. Log what is actually said about you on the prompt set, not just whether you appear.

Referral traffic and assisted conversions are the commercial end. Assistants increasingly link out, those visits arrive with intent, and they can be segmented in analytics like any other channel. The volumes are usually small, but they are real, and they anchor the programme to revenue rather than to counting.

Why a frozen prompt set is the whole game

Every metric above is a reading on a specific set of questions. Change the questions and you break the trend: the before and the after are no longer measurements of the same thing, and the month-over-month line, the only output a decision can rest on, becomes fiction. This is also how vanity dashboards get made. Rarely by lying; usually by quietly swapping hard questions for flattering ones until the chart only goes up.

So freeze the set. A few dozen questions your buyers genuinely ask is enough to be representative without becoming unmanageable. When the market moves, add new questions as a separately dated cohort with its own baseline; never silently replace the originals mid-series. If a question truly must go, retire it visibly and say so in the reporting.

Run daily, report monthly

The same prompt on the same engine can return a different answer an hour later, so any single reading is noise. The fix is cadence at two speeds. Run the prompt set frequently, daily where possible, and let the runs aggregate into a stable monthly figure. Then report monthly, against the baseline, on the unchanged set. The daily runs are for the operators; the monthly reading is for decisions, and it is also what a board or a client can actually absorb. We set out a one-page format for that audience in reporting AI visibility to your board. For the underlying observation method, meaning how to capture answers and citations in the first place, see how to measure AI search visibility.

Report per engine, never blended

ChatGPT, Gemini and Perplexity are not one channel. They retrieve from different sources, cite with different frequency, and answer the same buyer question differently; it is entirely normal to hold a strong citation rate on one and be absent from another. A blended figure hides exactly that, and it hides the work, because what improves selection in one engine often does little in the next. Report a small table per engine: citation rate, share of citations and accuracy, each with its trend since baseline. Three honest lines per engine beat one impressive number that answers to nothing.

Where Morris McLane fits

Measurement on these terms is a standing programme, not a one-off audit. Morris McLane runs it as a managed AI visibility monitoring service within our AI search visibility work: a fixed prompt set built from the questions your buyers actually ask, run per engine on a standing cadence, and reported month over month against the baseline — for organisations directly, and white-label for communications and government-relations firms reporting to their clients. If you want the metrics above on your own organisation, start here.

Frequently asked questions

How can I avoid vanity metrics in AI visibility tracking?

Apply three tests to any number before it goes in a report. Is it measured on a fixed set of questions your buyers actually ask, so it cannot be flattered by changing the prompts? Is it comparable month over month, with the same questions, engines and method each cycle? And does it connect to something commercial, such as being cited during buying research or traffic from AI surfaces? Raw mention counts, sentiment without citations, one-off screenshots and undisclosed blended scores all fail at least one test.

How do we measure AI search visibility month over month?

Freeze a set of the questions your buyers ask AI assistants and run them through each engine on a frequent, ideally daily, cadence. Record whether your domain is cited, who else is, and whether what is said about you is accurate. Aggregate the runs into a monthly reading per engine and compare it with the previous month on the same prompt set. Because the questions never change, any movement is real change in visibility, not an artefact of measurement.

Which AI visibility metrics actually matter?

Five. Citation rate: on a fixed set of buyer questions, how often your domain is cited in the answer. Share of citations: your citations as a proportion of all citations on those prompts, against competitors. Accuracy: whether what the engines say about you is correct. Referral traffic and assisted conversions from AI surfaces. And the month-over-month trend on an unchanged prompt set, which is what turns the other four into evidence of progress.

Why does the prompt set have to stay fixed?

Because the trend is the product. Every metric in an AI visibility programme is a reading on a specific set of questions; change the questions and the readings before and after are no longer comparable, so the trend line breaks. Swapping in flattering prompts is how vanity dashboards are made. Add new questions as a separately tracked cohort if the market moves, but never silently replace the originals mid-series.

Should AI visibility be reported per engine or as one blended number?

Per engine. ChatGPT, Gemini and Perplexity retrieve sources differently, cite with different frequency and answer the same question differently, so a blended figure can hide the fact that you are strong in one engine and invisible in another. The work needed to improve each engine differs too. Report a small per-engine table: citation rate, share of citations and accuracy for each, with the trend since baseline.

Are composite AI visibility scores reliable?

Treat any single blended score with caution unless the methodology is fully disclosed: which prompts, which engines, what weighting, and what counts as a mention. A score you cannot decompose cannot tell you what to fix, and it can rise or fall on a definition change rather than a real one. The underlying observations, meaning citations, share and accuracy on a fixed prompt set, are the reliable layer; a composite is at most a summary of them.

Related service AI Search Visibility Explore

More in AI Search Visibility

Reporting AI search visibility to your board or members

How a coalition or association lead should report AI search visibility to a board or membership — the four things worth measuring, a one-page format a non-specialist can read, and why baselines and trends beat vanity dashboards.

AI Search Visibility |Industry ·7 min read

Get in touch

Tell us a little about the situation — narrative, exposure, timing. We'll reply promptly with initial thoughts and next steps. Confidential, always.