You measure ChatGPT visibility by treating it as a sampling problem. For each prompt a buyer would realistically type, you collect many answers under controlled conditions, classify each one by hand or by a reliable classifier, and report rates: how often you are named, where you sit when named, how often your domain is linked, and how much of the total brand airtime you hold against competitors. One answer is an anecdote. A hundred answers is a measurement. This guide gives you the metrics, the sample sizes, the controls and a template you can run yourself or demand from any vendor.
Why a screenshot is not a measurement
ChatGPT does not have a fixed ranking to look up. Each answer is generated fresh, and the model samples its words, so the same prompt can name your brand in one run and omit it in the next. Live web retrieval adds another source of variation, because the pages pulled into the answer change from day to day. Personalization adds a third: a logged-in user with memory enabled sees answers shaped by their history.
So “ChatGPT recommends us” is only meaningful as a rate under stated conditions. “Named in 68 of 100 logged-out mobile answers from the US for the prompt ‘best expense management software for a 50-person company’, average position 2.1” is a measurement. “I asked it and we came up” is not. Everything else in this guide follows from taking that seriously.
The four metrics that describe visibility
| Metric | What you record per answer | What you report per prompt | What it tells you |
|---|---|---|---|
| Mention rate | Brand named: yes or no | Share of answers naming you | Whether you are in the consideration set |
| Position | Rank of your brand in the list, when named | Average position and share in top three | Whether you are a real candidate or an afterthought |
| Citation rate | Link to your domain present: yes or no | Share of answers linking to you | Whether retrieval grounds on your pages; the direct-click channel |
| Share of voice | Every brand named | Your mentions divided by all brand mentions | Your standing against the competitive set |
Record the competitor names in every answer even if you only care about yourself. The competitor set is often the most useful output of the whole exercise, because it tells you who the model considers your real peers, which may differ from who you think they are.
Mention rate
Count a mention only when the brand is clearly identified as a recommendation or option, not when it appears as a negative example (“unlike X, which lacks…”) or a passing reference. Decide the rule before you start and apply it identically to competitors.
Position
Position is the order in which brands are presented. In list-style answers it is unambiguous. In prose answers, use order of first appearance. When ChatGPT groups brands by sub-category (“for small teams: A, B; for enterprises: C”), record position within the answer as a whole and note the grouping; buyers read top to bottom regardless of headings.
Citation rate
Record a citation only for links to your own domain. A link to a review of you on a third-party site is useful context but is a different thing. Expect citation rate to be far lower and far noisier than mention rate. In 2026 research on AI citations, only 10.6% of cited URLs persisted across 28 days and 40 to 60% of cited sources rotated monthly. You are measuring a moving target by design.
Share of voice
If an answer names four brands and you are one of them, you hold 25% of that answer’s voice. Averaged across a prompt’s answers, share of voice shows whether the gap between you and the leader is closing even when your own mention rate is flat.
Build the prompt set before you measure anything
A measurement is only as good as the prompts. The common error is to test the prompts you wish buyers typed (“best AI-powered revenue intelligence platform”) instead of the ones they do (“software that tells sales reps which deals are about to fall through”). Build a prompt cluster that reflects real buyer language across the stages of a decision.
Suppose a cluster of 40 prompts for a B2B payroll product. It might split into:
- Category discovery: “best payroll software for small businesses in the UK”
- Constraint-led: “payroll software that handles contractors in five countries under $10 per person”
- Comparison and switching: “alternatives to [incumbent] for a growing startup”
- Use-case: “which payroll tool integrates with Xero and handles pensions auto-enrolment?”
- Trust: “is [your brand] legit?” and “[your brand] reviews”
Weight the cluster toward the prompts with commercial intent, and include a few you expect to lose; a baseline that only contains your strongest prompts cannot show you where the growth is. The craft of writing these is covered in prompt clusters: how buyers actually phrase the question, and our prompt generator will draft five to start from for any category, location and buyer type.
Sample size: how many answers per prompt
The honest answer is “more than feels necessary”. With five answers, a brand named in three of them has a true mention rate somewhere between roughly 20% and 90%. With twenty, the band narrows but still cannot reliably distinguish 50% from 65%. Our own audits read about 100 real answers per prompt per market, because at that volume a change of ten or fifteen points between two audits is likely to be real rather than noise.
Whatever you choose, three rules hold:
- Use the same number for every prompt and every audit so results are comparable.
- Report a range or a confidence band, not a single decimal. “Mention rate 62%, likely between 52% and 72%” is more truthful than “62.0%”.
- Do not compare a 10-answer audit this month with a 50-answer audit next month and call the difference progress.
The full description of our sampling and classification approach is on the page describing our audit method.
Controls: making conditions identical between audits
Variance you cannot explain is variance you cannot act on. Fix these conditions and keep them fixed.
| Condition | Why it matters | Recommended setting |
|---|---|---|
| Login state | Logged-in accounts with memory produce personalized answers | Logged out, or a clean account with memory and history off |
| Location | Answers change for anything with a local or regulatory element | Collect from genuine connections in the target country |
| Device | Mobile and desktop can differ in retrieval and formatting | Pick one; we use mobile because that is where buyers most often ask |
| Model and mode | Different models and “search” vs default change results | Record which was used; use the default a buyer would get |
| Session | Follow-up questions inherit context | One prompt per fresh session |
| Time window | Retrieval shifts day to day | Collect each audit within a short window and note the dates |
Collecting through a data center IP, through an API rather than the consumer product, or through a logged-in account with a history of asking about your own brand will all give you a picture of something other than what a buyer sees. Our audits collect from genuine mobile connections in the target country for exactly this reason.
Classifying answers consistently
Each answer gets read and coded. For every answer, record: prompt, date, brand named (yes/no), position (integer or blank), link to brand domain (yes/no), and the full ordered list of brands named. Two people coding the same twenty answers should agree on nearly all of them; if they do not, your rules are ambiguous and need tightening before you scale.
Resist the temptation to code only your own brand. The competitor list is what lets you compute share of voice and, more importantly, lets you see the model’s view of your category: which brands are the “default” answers in position one, which are niche alternatives, and whether a newcomer is climbing. Those are strategic signals that no amount of staring at your own mention rate will reveal.
Reading the results
A baseline audit typically produces one of a few patterns, each with a different implication.
- High mention rate, poor position: you are a known alternative but not the default. The work is in how comparison sources rank you, not in whether they include you.
- Low mention rate for buyer language, high for your own vocabulary: a category framing problem. Sources describe you in your words, buyers ask in theirs.
- Named in default answers but absent when ChatGPT searches the web: your coverage is stale. Model knowledge carries you; recent pages do not. Research in 2026 found cited content had a median age of 62 days for Claude versus 130 days for Google, so this pattern is common for established brands.
- Named but never linked: the model knows you but retrieval does not land on your pages. Check what pages are being cited instead, and whether yours are crawlable and factual.
- Absent everywhere: usually an entity problem or a consensus gap. See why competitors get named and you do not.
Present per-prompt tables, not a single blended score. A blended “AI visibility score” hides exactly the prompt-level differences that tell you what to do.
Re-audit on a fixed cadence
The baseline is only useful if you repeat it. Run the identical prompt set, with identical controls and sample size, at day 30 and then monthly. Plot mention rate, top-three share and share of voice per prompt over time. Expect uneven movement: some prompts respond within weeks, others take months, and model updates can move everything at once in either direction.
A re-audit also catches regressions. Because cited sources rotate so heavily, a brand that leaned on one or two strong pages can lose them without noticing. Monthly measurement turns that from a surprise into a line on a chart.
Measuring the outcome, not just the answer
Visibility inside ChatGPT is the leading indicator. The business outcome shows up elsewhere, and mostly not as ChatGPT referral traffic. In the SE Ranking study of over 100,000 sites, AI platforms accounted for 0.32% of all website traffic in 2026, with ChatGPT holding 74.78% of AI referrals. Small, and growing fast, but small.
The larger effect is indirect. A 2026 Idea Grove survey found 45% of consumers immediately Google a brand recommended by AI and only 2% would buy from an unfamiliar brand on the AI recommendation alone. So pair the ChatGPT audit with two outcome metrics: branded search volume and direct or organic brand-term sessions, and the referral line tagged utm_source=chatgpt.com. How to set that up is in our ChatGPT attribution guide.
Doing it yourself, using a tool, or having it done
A small team can run a credible manual audit on ten to twenty prompts with twenty answers each in a couple of days, and we have written the step-by-step DIY audit for exactly that. Beyond that scale, collection needs automating, and the tradeoffs between tracking tools and a done-for-you audit are laid out in AI visibility tools vs a done-for-you audit.
Our own audits are the baseline of The Blue Ocean GPT service: about 100 answers per prompt per market, read and classified, reported per prompt so your team can open ChatGPT and reproduce any result. The same audit is re-run identically at day 30 and monthly. If you would like to see what the baseline looks like for your brand, start with what ChatGPT currently says about you.
What to do next
- Write 10 to 20 buyer prompts in buyer language and fix the conditions you will test under before collecting a single answer.
- Record brand, position, link and every competitor for each answer; report per-prompt rates with a range, not one blended score.
- Put the day-30 re-audit in the calendar now and pair it with branded search volume as the outcome metric.
Frequently asked questions
How many times should I ask ChatGPT the same prompt to measure visibility?
Enough to see a stable rate rather than a lucky run. Five answers can tell you whether you are in the conversation at all. Twenty gives you a rough rate. Our own audits read about 100 answers per prompt per market because that produces a confidence band narrow enough to compare month to month. Whatever number you choose, keep it constant between audits.
What is a good mention rate in ChatGPT?
It depends on the prompt and how many brands fit it. For a specific prompt with three or four obvious candidates, the leader is often named in most answers. For broad prompts with dozens of plausible brands, even 40% is strong. Judge yourself against the top competitor in your own audit, not against an absolute figure, and watch the trend across re-audits.
What is the difference between mention rate and citation rate?
Mention rate is the share of answers that name your brand in the text. Citation rate is the share of answers that include a link to your domain among the sources. Mentions reflect the model's accumulated view of your brand and are relatively stable. Citations reflect which pages were retrieved that day and rotate heavily, so they are a weaker signal of reputation but the only one that produces direct clicks.
Can I measure ChatGPT visibility with a tracking tool instead of manually?
Tools can automate the collection, which matters once you are tracking dozens of prompts monthly. The questions to ask any tool are how many answers it samples per prompt, whether it collects from real sessions in your target country on realistic devices, whether it classifies position and competitors, and whether you can reproduce a given result yourself in ChatGPT. Automation does not fix a bad prompt set.
Why do my ChatGPT results differ from what my colleague sees?
Because ChatGPT answers are personalized and probabilistic. A logged-in account with memory enabled, a different location, a different device, a different model version or simply a different run will produce different shortlists. That is why audits are done logged out or in clean sessions, from the target market, on consistent devices, and across many answers rather than one.
Find out what ChatGPT says about your brand before you spend anything.
Apply in about two minutes. We check your keywords, then run your real buyer prompts live on a 30-minute call. If we cannot create the result for your keywords, we tell you on the call, not after an invoice.
See if my keywords qualify →