Every visibility number on this site comes from one procedure: about 100 real ChatGPT answers per prompt per market, collected from genuine mobile connections inside the target country, each read by a person and classified against four questions. Is the brand named? Where in the list? Is the brand’s own domain linked? Which competitors are named? The audit is re-run identically at day 30 and monthly. This page explains each step, then spends equal space on what the method cannot tell you.
Why one answer from ChatGPT is not a measurement
Ask ChatGPT “best accounting software for a 20-person agency” ten times and you will usually get the same three or four brands, not always in the same order, and the fifth and sixth names will change. The model samples from a distribution. Sometimes it runs a live web search, sometimes it answers from model knowledge, and the two routes pull from different sources. The source layer is unstable: a Digital Authority Partners study found only 10.6% of cited URLs persisted across 28 days. One screenshot, good or bad, proves very little. What a buyer experiences is the distribution, and the audit’s job is to estimate it well enough to act on.
Step 1: Build the prompt cluster
We start from how buyers phrase the question, not from your Google keyword list. A prompt cluster for one commercial intent typically holds 20 to 40 prompts: direct asks (“best X for Y”), comparisons (“X vs Z for a small team”), situational asks (“I need an X that works with Y under Z dollars”), and local or trust asks (“is X legit”, “X in Austin”). The prompt cluster guide covers construction, and the free ChatGPT prompt generator tool produces a five-prompt starter set. The cluster is agreed before collection and does not change for the life of the engagement; a new intent becomes a second cluster with its own baseline audit.
Step 2: Collect about 100 real answers per prompt per market
For each prompt we collect roughly 100 answers in the target market, spread over several days, because a burst collected in one afternoon measures that afternoon, not the month. Three conditions are held constant:
- Genuine mobile connections in the target country. Not datacenter IPs or VPN exits; ChatGPT responds to where the request appears to come from.
- Logged-out sessions. Logged-in accounts carry memory and personalization that would contaminate the sample.
- The same prompt text, byte for byte. No paraphrasing between runs.
Weekly tracking runs between full audits use smaller samples (about 20 answers per prompt) to watch direction; only the 100-answer runs are used for before-and-after comparisons.
Step 3: Read and classify every answer by hand
Each answer is read by a person and recorded against four fields. We do not rely on keyword matching, because ChatGPT names brands in ways a regex misses: “the Shopify-owned option”, “the one most agencies default to”.
| Field | What is recorded | Why it is tracked separately |
|---|---|---|
| Mention | Brand named anywhere in the answer, yes or no | The base measure of being in the conversation |
| Position | Order in which the brand appears; first named is 1 | Buyers shortlist from the top; see why position 1 to 3 is the whole game |
| Citation | A link to the brand’s own domain present in the sources, yes or no | Being named and being linked are different things, explained in mentions versus citations |
| Competitors | Every other brand named, with its position | Gives share of voice and shows who you are displacing or losing to |
From these fields we compute, per prompt, the mention rate (answers naming the brand divided by total answers), the average and median position when named, the citation rate, and share of voice against every named competitor. Per cluster, we report the distribution, not only the mean.
Step 4: Re-run identically at day 30 and monthly
The re-audit uses the same prompts, market, connection type, classification rules and roughly the same sample size; nothing else is comparable. Movement is reported as a change in mention rate with its confidence band. A move from 12% to 18% on 100 answers each side is inside the band and we say so. A move from 12% to 55% is not.
Every report lists the prompts, so you can open ChatGPT and see for yourself; if you cannot recognize the pattern at the same prompt, the number should not be trusted. You can see how ChatGPT describes your brand today by asking us for a baseline.
What the method cannot tell you
Every limit below applies to our audit, and even more strongly to a casual check in a browser.
Answers vary day to day
Even with everything held constant, the same prompt yields different answers across days because the model samples and the search layer rotates. A brand at 60% one week can genuinely read 52% or 68% the next with nothing changed on its side. We call a change real only when it clears the band.
100 answers gives a confidence band, not a decimal
With 100 answers, a measured mention rate of 50% has a 95% confidence interval of roughly 40 to 60%. At 10% or 90% the band narrows to about plus or minus 6 points, but it is still a band. We print “58%” because tables need a number; the honest reading is “mid fifties to low sixties”. A mention rate with a decimal point from a sample this size is noise reported as precision. See sample size and variance in the glossary.
Logged-in and logged-out are different products
A logged-in account has memory, custom instructions and, on free tiers in the US, a Sponsored card under some answers. Our logged-out collection measures the shared baseline: the answer a new buyer with no history gets, which will not match your own account.
Mobile and desktop differ
Source panels, the number of brands shown and whether a search is triggered can differ between the mobile app and the desktop client. We collect on mobile because that is where most consumer and many B2B buyers ask; if yours are mostly on desktop, say so and we adjust.
Location changes the answer
“Best immigration lawyer” asked from Toronto and from London produces different lists, because the search layer and the model’s priors both respond to location. A market in our audit is one country. Three countries means three audits, never averaged into one headline number.
What a mention rate does not measure
A high mention rate is not revenue. For nearly half of consumers the next step is a Google search for your brand: a 2026 Idea Grove survey found only 2% would buy from an unfamiliar brand on an AI recommendation alone and 45% immediately Google the brand. So watch brand search volume and utm_source=chatgpt.com referrals alongside the audit, as described in how to track ChatGPT traffic. The audit says whether you are recommended; analytics says whether it pays.
What we do not report
We do not report prompts outside the agreed cluster, present competitor rates as our client’s, or extrapolate from one market to another. We also do not describe how the mention rate is moved; how we do it is explained on your call.
How to read the report
Each monthly report has three layers: a cluster summary (average mention rate, share of prompts with a top-three position, share of voice against the five most-named competitors, each with its change from baseline and whether it clears the band), a prompt table (mention rate, median position and citation rate per prompt), and a competitor view.
What to do next
- Generate five prompts with the prompt generator tool and ask each three times in a logged-out ChatGPT session on your phone.
- Record mention, position, citation and competitors for each answer, following the step-by-step DIY audit guide. Fifteen answers show the pattern, not a trustworthy percentage.
- For the full 100-answer baseline across a real cluster, request it from the homepage and we walk you through the results prompt by prompt.
Frequently asked questions
Why do you collect 100 answers per prompt instead of asking ChatGPT once?
Because ChatGPT does not give the same answer twice. The same prompt, asked an hour apart, can name a different set of brands in a different order. One answer is an anecdote. A hundred answers let us say "named in roughly 60 to 70 percent of answers" with a confidence band we can defend, and let us see movement at day 30 that is larger than the noise.
Can I reproduce your numbers myself in ChatGPT?
Yes, at the prompt level. Every prompt we test is listed in your report with its mention rate, average position and citation rate. Open ChatGPT, type the prompt, and you will see the kind of answer we classified. You will not match our percentages exactly from a handful of tries, because of variance, but you should see the same brands recurring in the same rough order.
Does it matter whether I am logged in, on mobile, or in another country?
Yes, all three change answers. Logged-in accounts carry memory and personalization, mobile and desktop can show different source panels, and location changes which local brands appear. We collect from logged-out sessions on genuine mobile connections inside the target country so that every audit measures the same thing. Your own checks will differ if you change any of those conditions.
What counts as a mention, a position and a citation in your audit?
A mention is the brand named anywhere in the answer. Position is the brand's place in the list or ordering ChatGPT gives; first named is position one. A citation is a link to the brand's own domain present in the answer's sources. These are tracked separately because a brand can be named without being linked, and linked without being recommended.
How often do you re-run the audit?
Identically at day 30 and then monthly. Same prompts, same market, same connection type, same classification rules. Changing any of those between runs would make the comparison meaningless, so the baseline audit fixes them for the whole engagement.
Find out what ChatGPT says about your brand before you spend anything.
Apply in about two minutes. We check your keywords, then run your real buyer prompts live on a 30-minute call. If we cannot create the result for your keywords, we tell you on the call, not after an invoice.
See if my keywords qualify →