Gemini names a handful of brands in every answer, then rebuilds that list the next time someone asks. Here is the method that turns those moving answers into a number you can track.
Run CITED's free AI visibility check to see where you stand in Gemini right now, or go straight to the Gemini tracker if you already know you need this measured daily.
Ask Gemini the same fixed list of buyer questions on a set schedule, record whether your brand is named in each answer, and score the results as one number: the percentage of prompts that mention you. A single answer is a sample, not a measurement. The trend across repeated runs is the measurement.
Before you can count anything you have to decide what a hit is, and this is where most tracking spreadsheets go wrong on day one. Gemini can involve your brand in an answer in five distinct ways, and they are not worth the same amount.
The distinction that matters most is between a mention and a citation. A mention is your brand name appearing in the text Gemini wrote. A citation is your domain appearing in the grounding links underneath it. You can get one without the other in both directions, and the fix for each is different, so count them in separate columns.
| Outcome | What it looks like | What to do with it |
|---|---|---|
| Named | Gemini writes your brand name somewhere in the answer | Your baseline hit. This is what a visibility score counts |
| Recommended | Gemini names you as a pick and gives a reason why | The outcome that actually moves revenue. Worth its own column |
| Cited | Your domain appears in the sources under the answer | A citation, not a mention. Track it separately: it is the earliest signal that your content is being read |
| Generic answer | Gemini answers the question without naming any brand | An open prompt. Nobody owns it yet, so it is the cheapest one to win |
| Competitor only | Gemini names rivals and not you | Your clearest content brief. Read the answer and write the page it wanted |
This is the single most common measurement error, and it is expensive because it feels like diligence. Someone checks the Gemini app, sees the brand named, and reports that the company is visible in Google's AI. Meanwhile the AI Overview on the same query has been recommending a competitor for a month.
All three run on Gemini models. That is where the similarity ends. They retrieve different documents, they are triggered by different inputs, and they answer to different audiences. Treat them as one metric and you will average away the surface that is actually costing you customers.
| Surface | Where it appears | Why the answer differs |
|---|---|---|
| Gemini app | gemini.google.com and the mobile apps | A conversation. Answers carry context from earlier turns and from the account asking |
| AI Overviews | The top of an ordinary Google results page | Triggered by a typed search, not a question. Shorter, more cautious, and tuned to the query's intent. See the AI Overview tracker |
| AI Mode | The AI tab inside Google Search | Fans one question into many background searches, then merges them. Names more brands, and different ones. See the AI Mode tracker |
Gemini does not hold a ranked list of brands in memory and read it out. For most commercial questions it runs a search first, reads what comes back, and writes an answer grounded in those documents. That grounding step is the reason your mentions move, and it is also the reason they are winnable.
Four things change between two runs of the same prompt. Knowing which one moved is the difference between a real trend and a false alarm.
Gemini grounds its answer in whatever the search step returned this time. A new article, a reordered results page, or a competitor's fresh comparison post changes the raw material before the model writes a word.
"Best project management tools" and "best project management software for small teams" are different retrieval problems and often produce a different brand list. Retyping a prompt from memory instead of pasting it is enough to break a time series.
Country, language, and account history all feed the answer. A check run from a US account and a check run from a German one are two separate measurements, and mixing them into one number tells you nothing about either market.
Google ships model updates continuously and rarely announces what moved. A step change across every prompt on the same day, for you and your competitors alike, is usually this rather than anything you did.
You do not need a tool to start, and starting by hand is worth doing once even if you automate later. It teaches you what your buyers actually ask and what Gemini actually says back, which no dashboard can hand you second-hand.
Not keywords. Questions, phrased the way a customer would type them. Pull them from sales calls, support tickets, and your own search data. Cover the funnel: the category question, the comparison, the alternatives-to, and the specific use case. Ten good prompts beat a hundred padded ones.
Same account or a consistently signed-out session, same country, same language, same time window. Paste the prompt from your list rather than retyping it. Start each prompt in a fresh chat so the previous answer cannot bias the next one.
One row per prompt per run: date, the exact prompt, yes or no on your brand, which competitors were named and in what order, and the source links shown. Paste the answer text too. The wording is where the content brief hides.
Count the prompts that named you, divide by the prompts you ran, and write down the percentage. Seven hits out of ten prompts is 70 percent visibility for that run. One number per run is what makes two runs comparable.
Weekly is the practical floor for a manual process. Never act on the gap between two consecutive runs: a prompt either names you or it does not, so single checks swing hard by design. Four or five runs is when the line starts meaning something.
A spreadsheet of answers is not a measurement until it collapses into numbers you can put in front of someone. These four are the ones worth reporting, and they are the ones CITED computes on every check.
Note what is deliberately missing: a volatility percentage. You will see articles quote a precise figure for how much AI answers swing. Those numbers are unsourced, and they vary so much by category that a single figure would be misleading. The honest version is the mechanism above: each prompt is a yes or no, so read the trend line, not the last point.
| Metric | What it measures | How to read it |
|---|---|---|
| Visibility | The share of your tracked prompts where Gemini names your brand | The headline number. 70% means named in 7 of 10 answers. Compare it to itself over time, not to an industry benchmark |
| Share of voice | Your mentions as a percentage of every brand mention across the whole prompt set | Corrects for crowded answers. Visibility can hold at 70% while your share falls because Gemini started naming eight brands instead of four |
| Position | Your average place in the list when you are named | #1.6 means you are usually first or second. Being named last in a list of nine is not the same win as leading it |
| Citations | How often your own domain appears in the sources Gemini grounded on | The leading indicator. Your domain getting retrieved usually moves before your mention rate does |
A 40 percent visibility score sounds like a failure until you learn the category leader sits at 45 and everyone else is under 20. It sounds like a triumph until you learn the leader is at 90. The number on its own is not information.
So log every brand Gemini names, not just your own. It costs nothing extra while you are reading the answer anyway, and it converts your score from a lonely percentage into a ranking. It also surfaces the competitors you did not know you had: the ones Gemini treats as your alternative are frequently not the ones on your sales battlecard.
An overall visibility score averages two very different problems together. A brand at 70 percent might be winning seven prompts consistently and losing three consistently, or it might be winning all ten about two thirds of the time. Those need opposite responses, and the headline number cannot tell them apart.
Split your prompts into three buckets instead. The ones you always win need defending. The contested ones respond fastest to work, because Gemini already considers you a candidate. The ones you never win are where the real gap is, and they are the honest content roadmap: each one is a question your buyers ask that your site has never answered well enough to be retrieved.
Because Gemini grounds its answers in retrieved documents, the brand list in any answer is downstream of a more useful question: which pages did it read to get there? Tracking only the mention tells you that you lost. Tracking the sources tells you where you lost.
In practice a handful of domains do most of the work in any category, and they are rarely the ones you would guess. Review sites, comparison posts, and a couple of forum threads usually outrank vendor pages as source material. Logging which of them fed each answer, and whether your brand was named on them, converts a visibility problem into a specific list of pages to go and get mentioned on.
Any guide that does not have this section is selling you something. Four things are genuinely out of reach, and knowing them keeps you from over-reading your own data.
Volume. There is no impressions figure. You cannot learn how many people asked Gemini your tracked question last month, so visibility is a rate, never a traffic estimate.
Attribution. Traffic arriving from a Gemini answer often carries no referrer and lands in analytics as direct. You can watch branded search and direct traffic move together with your mention rate, but you cannot close the loop cleanly.
Why. Gemini does not explain why it picked three brands out of thirty. The grounding sources are the closest thing to a reason, and they are a strong clue rather than an answer.
Everyone else's answer. Your measurement describes the conditions you fixed: your region, your language, your account state. It is a reliable index of your own direction over time, not a census of what every user sees.
Do it by hand until the arithmetic turns against you. Twenty prompts logged properly takes something like an hour, so a weekly manual check is a few hours a month and entirely reasonable. The trouble is that the method's accuracy depends on exactly the parts people quietly drop when they are busy: identical wording, a clean session, the same day of the week, every competitor recorded rather than just the ones you recognised.
Automate at the point where the log stops being trustworthy, which in practice is around 50 prompts or the first missed week. A tracker also buys you a frequency a person cannot sustain: CITED re-asks your prompt set every day, holds the wording and the region fixed by construction, records every brand and source automatically, and turns the losing prompts into a ranked list of things to fix rather than a spreadsheet you still have to interpret.
The Gemini app passed one billion monthly users in August 2026, making it the fastest-growing product Google has ever shipped. Every one of those conversations can name a short list of brands, and none of them appear in a rank tracker.
Yes. A tracker re-asks a fixed prompt set on a schedule and scores each answer for you, which removes the two things that break manual logs: drifting prompt wording and missed runs. CITED checks daily and holds the prompt set frozen so every check stays comparable to the last.
Ten is enough to get a real baseline, and 50 is enough to cover a category properly. Past that you are usually adding paraphrases rather than new questions. Prompt quality matters far more than count: ten questions your buyers genuinely ask beat a hundred keyword variations.
Daily if it is automated, weekly if you are doing it by hand. The point of the frequency is not to react to each check, it is to build enough points that the trend line is readable. Single checks swing because each prompt is a yes or no.
A mention is your brand name appearing in the answer Gemini wrote. A citation is your domain appearing in the grounding sources listed underneath it. You can be cited without being mentioned, and mentioned without being cited. Track them in separate columns, because citations usually move first.
Ranking and being named are different mechanics. Gemini reads a set of retrieved pages and writes a short list from what those pages say. If the pages it reads discuss your competitors and not you, a first-place ranking on your own site does not put you in the answer.
Yes, and noticeably. Region and language change which pages are retrieved, which changes the brand list. Track one market per prompt set rather than mixing them, otherwise a change in your score could just be a change in where you happened to run the check from.
Yes, and you should use the same prompt list for both so the results are comparable. CITED tracks ChatGPT, Gemini, and Perplexity together. Our guide to monitoring brand mentions across every LLM covers the cross-engine workflow in full.
Not for manual tracking, which only needs the Gemini app and a spreadsheet. An API with search grounding is how automated trackers run checks at scale and capture the source links programmatically, but that is the tool's problem to solve, not a prerequisite for you to start measuring.
Run a free instant check on your own buyer questions. No login, no credit card, about 15 seconds.
Check my Gemini visibility →