How to measure AI visibility in Gemini
There is no Gemini ranking to look up. Google publishes no position for any of its AI surfaces, so visibility here is measured the same way it is measured in every generative engine: by repeated observation against a baseline. What makes Gemini different, and harder than the others, is that “Gemini” is not one thing. Before any metric means anything, you have to decide which surface you are measuring.
This is the per-engine layer beneath how to measure AI search visibility. The method there applies here; what follows is what changes when the engine is Google’s.
The three surfaces people report as one
Google’s AI can represent your organisation in at least three distinct places, and they are not interchangeable.
- AI Overviews. The generated summary that appears above or among conventional Google results. It sits inside Search, it is triggered on a subset of queries, and it typically shows links to the sources it drew from.
- AI Mode. Google’s conversational mode inside Search, where a query becomes a dialogue rather than a page of results. Retrieval is search-led, and answers carry sources.
- The Gemini assistant. The separate product at its own app and site. Some answers are grounded in a live search and carry sources; others are answered from the model without them.
They share infrastructure and they overlap, but they do not behave alike. Whether a given answer is grounded in a live search varies by prompt and is not guaranteed, which means the same question can produce a well-sourced answer in one surface and an unsourced one in another. So it is entirely normal to hold a decent citation rate in AI Overviews while being absent from the assistant. A single blended “Gemini visibility” figure hides exactly that, and with it hides the work, because what improves selection in a search-led surface is not identical to what improves an ungrounded model answer.
Report three lines, not one.
The metrics that transfer, and the one that does not
Four of the five metrics in our AI visibility metrics guide carry over to Gemini without modification: citation rate on a fixed prompt set, share of citations against competitors on those prompts, accuracy of what is said, and the month-over-month trend on an unchanged set.
The one that needs care is citation rate, because it silently assumes links are shown. Where a surface does not cite, a citation rate of zero is not evidence of invisibility; it may only be evidence that this answer was not grounded. That distinction matters commercially, so log two things separately:
- Whether you were mentioned at all, in the text of the answer, links or no links.
- Whether you were cited, meaning your domain appeared among the sources.
Being described accurately in an answer that links to nobody is still visibility. It is worth less than a citation, because it sends no traffic and offers the reader no way to verify, but recording it as a zero throws away a real reading. Where sources are shown, citation rate is the stronger signal and should carry the weight.
Building a Gemini prompt set
The set is the measurement instrument, so it is built once and then left alone.
Write the questions your buyers actually ask. Commercial questions that precede a shortlist, not brand searches. If someone already knows your name, their query tells you little about whether the engine would have surfaced you.
Hold the set fixed across all three surfaces. The same questions, run everywhere, so the surfaces are comparable with each other as well as with their own baselines. This is what lets you say “we are cited in AI Overviews and absent from the assistant” and mean something by it.
Fix the variables you can. Location and language change answers, so pin them and record what you pinned. A reading taken from a different country than your buyers is not a reading about your buyers.
Expect the models to move underneath you. Google updates on its own schedule and gives you no changelog for your prompt set. This is not a reason to avoid measuring; it is the reason to keep the prompts frozen, so that when a reading jumps you can at least tell it apart from a change you made.
Why sampling discipline matters more here
Every generative engine is noisy between runs. Google’s surfaces are noisier than most, because grounding is variable, personalisation is real, and there are three surfaces to keep straight. The consequence is practical: the single check is actively misleading in this engine, more so than in others.
Run frequently, daily where you can, and let the runs aggregate into a monthly figure per surface. Report that figure against the baseline. Any Gemini number presented without its prompt set, its surface and its run count cannot be verified and should not be acted on, including numbers presented by a vendor.
What measurement will not do
It will not change what Google says about you. There is no console where an organisation edits its own representation in an AI surface, and no markup that admits you to one; Google’s published position is that normal Search eligibility is what applies. Answers are assembled from sources, so the work is at the source layer: the pages you control made explicit and quotable, and the third-party sources an answer is actually being built from.
What measurement does is make that work honest. It tells you which surface you are absent from, which sources each answer was built from, and whether a fix propagated. The correction sequence itself is set out in how to fix AI misinformation about your brand.
Where Morris McLane fits
We run this as a standing programme rather than a one-off reading: a fixed prompt set built from the questions your buyers ask, measured per Google surface on a regular cadence, and reported month over month against the baseline, as part of our managed AI visibility monitoring within AI search visibility work. It runs for organisations directly, and white-label for communications and government-relations firms reporting to their own clients.
If you want a first read on where you stand across all three surfaces before committing to a cadence, an AI visibility audit is the one-off version of that baseline, run as the gap analysis described here.
Frequently asked questions
What metrics should I look at inside a Gemini visibility tracker to know if my AI SEO or generative engine optimisation efforts are working?
Four, each read per surface rather than blended. Citation rate: on a fixed set of your buyers' questions, how often your domain appears among the sources shown. Share of citations: your citations as a proportion of all citations on those same prompts, against the competitors that matter. Accuracy: whether what is said about you is correct and current, logged whether or not a link appears. And the trend on an unchanged prompt set, which is the only one of the four that tells you the work is landing rather than that the weather changed. Read them separately for AI Overviews, AI Mode and the Gemini assistant, because a gain in one does not imply a gain in the others.
How do you track visibility in Gemini?
By repeated observation against a baseline, because Google publishes no ranking for any of its AI surfaces. Freeze a set of the questions your buyers actually ask, run them through each Google surface on a standing cadence, and record for every run whether you were cited, who else was, and whether the description of you was accurate. Aggregate the runs into a monthly reading per surface and compare it with the baseline on the same unchanged set. A single check tells you almost nothing, because the same prompt can answer differently an hour later.
Is appearing in Google's AI Overviews the same as appearing in Gemini?
No, and conflating them is the most common measurement error in this engine. AI Overviews and AI Mode sit inside Google Search; the Gemini assistant is a separate product reached through its own app and site. They draw on overlapping but different retrieval paths and they cite with different frequency, so it is entirely normal to be well represented in AI Overviews and absent from the assistant, or the reverse. Measure and report each surface on its own line.
Why do Gemini answers change between runs of the same prompt?
Several reasons at once, which is why the sampling discipline matters more here than elsewhere. Whether the model grounds an answer in a live search varies by prompt and is not guaranteed. Answers can differ by location, by language and by the state of the account asking. And the underlying models are updated on Google's schedule, not yours. None of that makes measurement impossible, but it does mean any single screenshot is one run of one surface on one day, and should never be reported as a finding.
Can you measure Gemini visibility manually, or do you need a tool?
You can start manually, and for a small prompt set that is a reasonable baseline: run your priority questions through each Google surface, capture the full answer and any sources shown, and log presence, accuracy and the cited domains in a spreadsheet. The limitation is not accuracy but repeatability, because the value of this measurement comes from frequent runs aggregated over time, and that becomes unmanageable by hand quickly. The discipline matters more than the tooling either way: a fixed prompt set checked by hand beats a dashboard pointed at prompts nobody asks.
How do you correct wrong information about your company in Gemini?
Not in Gemini, which is the point people find frustrating. There is no console where an organisation edits what a Google AI surface says about it. Answers are assembled from sources, so correction is source-layer work: fix the authoritative pages you control, make the correct facts explicit and quotable rather than implied, and address the third-party sources an answer is actually being built from. Then re-measure on the same prompts to see whether the change propagated, because it does not happen on submission and it does not happen instantly.
Are AI visibility metrics for Gemini accurate?
They are accurate as observations and unreliable as single readings, and the difference is the method. Any one run is a genuine record of what that surface said at that moment, but the variance between runs is high enough that one run cannot support a decision. Aggregated over a frequent cadence on an unchanged prompt set, the same observations become a stable and comparable figure. Treat any Gemini number quoted without its prompt set, its surface and its run count as unverifiable.