There are four ways to measure how often ChatGPT recommends you, and the right one depends on whether you need a dashboard to act on, a number to report, or a result to show. This page compares the categories, not named products, so you can pick the one that fits how your team works.
The direct answer: pick by who does the work afterward
A tracking tool is right when someone on your team will read the data every week and change what the company publishes, pitches and fixes as a result. An enterprise suite is right when that someone is a team across several brands and markets. A spreadsheet is right when you need a first baseline this week with no budget line. A done-for-you audit is right when you want the measurement and the movement from one accountable party.
Every credible approach does the same thing: builds a prompt cluster of the ways buyers phrase the question, collects many ChatGPT answers per prompt, and classifies each answer for whether the brand is named, in what position, with what link, and alongside which competitors. The differences are in sampling depth, who classifies, and what happens next.
One caveat applies to all four: ChatGPT’s cited sources are unstable. A Digital Authority Partners study of 1,127 cited URLs between November 2025 and February 2026 found only 10.6% persisted across 28 days, and 40 to 60% of cited sources rotate monthly. Sampling once per prompt measures weather, not climate.
The four options side by side
| Dimension | Prompt-tracking dashboard | Enterprise AI visibility suite | Manual spreadsheet audit | Done-for-you audit plus service |
|---|---|---|---|---|
| What you get | Charts of mention rate, position, competitors, cited sources over time | The same, plus multi-brand, multi-market, team workflows and integrations | A sheet you built, with whatever you recorded | A baseline report, monthly re-audit, and the work to move the numbers |
| Who does the measuring | Software, on a schedule | Software, on a schedule | You | The provider, with a documented method you can reproduce |
| Sampling depth | Varies by product and plan; check how many runs per prompt | Usually configurable; higher plans sample more | As many as you have patience for | About 100 answers per prompt per market in our method |
| Who acts on it | Your team | Your team, often with vendor support | Your team | The provider |
| Cost drivers | Prompts tracked, engines, sampling frequency, seats | Brands, markets, users, data retention, integrations | Your hours | Clusters, markets, scope of the engagement |
| Time to first data | Hours to a day | Days to weeks (onboarding) | Same day | Days for the baseline |
| Best for | Marketers who will act weekly on the data | Multi-brand or multi-market organizations | First baseline on a tight budget | Teams that want the outcome, not another dashboard |
| Main risk | Data nobody acts on | Cost and complexity outrunning the use | Inconsistent sampling, small samples, no history | Choosing a provider who will not show the method |
Prompt-tracking dashboards: when they win
These are the SaaS products that do for AI answers what rank trackers did for Google: enter prompts, pick engines, get recurring charts.
They win when you have an in-house person who owns AI visibility and will treat the dashboard as a weekly to-do list rather than a report. The useful loop looks like this: the dashboard shows your mention rate on “best accounting software for agencies” prompts slipped from 40% to 25% this month; you look at which sources the answers now cite; you notice a new roundup that excludes you; you reach out to its author with a correction and a data point. Without someone to close that loop, the dashboard is a subscription to bad news.
Check before buying:
- How many times is each prompt run per period? One run is noise.
- From where? ChatGPT answers differ by country and by logged-in state.
- Can you see the raw answers, or only the aggregate? You need the raw text to understand why the number moved.
- Does it classify position in the answer, or only presence?
- Can you export the data so you are not locked in?
For what the charts should contain, read how to measure your brand’s visibility in ChatGPT.
Enterprise AI visibility suites: when they win
The enterprise tier adds many brands, markets and languages, role-based access, analytics integrations, long data retention and a customer success contact.
They win when the coordination problem is bigger than the measurement problem. A group with eight consumer brands in twelve markets cannot run eight dashboards and a shared spreadsheet; it needs one source of truth that regional teams can filter. It also wins when procurement requires security reviews, SSO and contracts.
The main cost driver is not the software but the organization around it: the people who configure prompt clusters per market, maintain competitor lists, interpret the data and brief the content, PR and product teams. Without that staffing, the suite is an expensive dashboard nobody reads. Our comparison of GEO agency vs in-house vs done-for-you covers the staffing question in more detail.
Manual spreadsheet audits: when they win
The spreadsheet audit is the zero-software option: prompts in column A, one row per run, columns for named yes/no, position, link present, competitors named, and notes on what the answer said.
It wins when you need a first baseline fast. Reading forty ChatGPT answers about your category by hand teaches you more about how the model describes you (and misdescribes you) than any chart. It also wins for very small clusters, say ten prompts for a local service business.
Its weaknesses are sample size and consistency. Most people run each prompt once or twice, which cannot distinguish a 20% mention rate from a 40% one. Five to ten runs per prompt is a workable DIY baseline; 20 is the minimum before you act on a figure. And a second audit a month later, by a different person in a different browser state, is rarely comparable. If you go this route, follow the step-by-step method in how to run a ChatGPT visibility audit yourself, use the prompt generator to build a realistic cluster, and write down your sampling rules before you start so the re-audit matches.
Done-for-you audit plus service: when it wins
This category bundles the measurement with the work: the provider builds the cluster, collects the baseline, works to move the numbers, and re-measures with the same method at day 30 and monthly.
It wins when the business wants an outcome rather than an instrument. The question a founder asks is not “what is our mention rate?” but “why is our competitor named in 70% of answers and we are named in 10%, and what will it take to flip that?” A dashboard answers the first question. A done-for-you service is accountable for the second.
Our own version, so you can judge any provider against it: the audit collects about 100 real ChatGPT answers per prompt per market from genuine mobile connections in the target country, and each answer is read and classified for brand named, position in the list, link to the brand’s domain, and every competitor named. The identical run is repeated at day 30 and monthly. The service targets first movement within the first week, consistent mentions driven toward up to 90% of prompts in a cluster, and position in the top three. Those are targets, not guarantees, and reporting is at prompt level so you can reproduce it in ChatGPT yourself. How we do it is explained on your call. The full method, including its limits, is on the methodology page.
What to demand from any provider in this category:
- A written sampling method (how many answers, from where, logged in or out, mobile or desktop).
- Prompt-level results you can reproduce in ChatGPT yourself.
- A re-audit on the identical method, not a new one that flatters the result.
- No lock-in, so the re-audit is the thing that keeps you, not the contract.
- Nothing that involves buying mentions, fake reviews or paid placements dressed as earned coverage. See buying brand mentions vs earning the recommendation for why those underperform anyway.
When to combine them
The most common mature setup uses two of the four.
Spreadsheet first, then a tool. Run a manual audit of 15 to 20 prompts to learn how ChatGPT talks about your category. Then, once you know which prompts matter, put those into a dashboard so the tracking is consistent and the history accumulates.
Dashboard plus done-for-you. An in-house marketer keeps a dashboard for daily awareness while a done-for-you provider runs the baseline and the work on the priority clusters. The dashboard becomes an independent check on the provider’s reporting. Expect the levels to differ because of sampling; the trends should agree.
Which to fund first: a decision rule
Answer two questions.
Do you already know whether you have a problem? If not, spend one afternoon on a spreadsheet audit or, for a measured answer across a full cluster, see what ChatGPT says about your brand.
If you have a problem, who will fix it? If a named person on your team will own it weekly, buy a dashboard (or a suite if you are multi-brand) and give them the time. If no one can own it, or the gap to the competitor is large and the category is one where the recommendation decides the shortlist, fund the done-for-you option and use the monthly re-audit as your accountability.
The rule compressed: measure cheaply first, then pay for whichever option matches who will do the work.
What to do next
- Write down your top twenty buyer prompts and run five of them in ChatGPT today; note who gets named and where you appear.
- Decide honestly whether anyone on your team has four hours a week for this. The answer determines the category.
- Read prompt clusters before you configure any tool, because a badly built cluster produces confident, useless numbers in all four options.
Frequently asked questions
What do AI visibility tools actually measure?
Most track a set of prompts across one or more AI engines and record whether your brand is named, its position, which competitors appear and which sources are cited. The better ones sample each prompt several times, because answers vary between runs. The number you want is mention rate (share of answers naming you) per prompt cluster, plus position, tracked over time and against competitors.
Can I measure ChatGPT visibility without paying for a tool?
Yes. Build a list of buyer prompts, ask each one in ChatGPT several times (ideally in a fresh session, from the target market), and record in a spreadsheet whether you were named, where, and who else was. It takes hours rather than minutes and gives a confidence band rather than a decimal, but for a first baseline on ten to twenty prompts it is entirely workable.
Why do AI visibility tools give different numbers for the same brand?
Because ChatGPT does not give the same answer twice. Which sources it retrieves, whether it uses live search, the user's location, device and login state all change the output, and 40 to 60% of cited sources rotate month to month. Two tools sampling at different times, from different locations, with different sample sizes will legitimately disagree. Compare trends within one method, not single numbers across methods.
Is a done-for-you audit just a tool with a human attached?
Partly, but the difference that matters is what happens after the number. A dashboard tells you that you are named in 15% of answers. A done-for-you service measures that same baseline, then does the work intended to move it, and re-measures with the identical method each month. You are paying for the outcome and the accountability, not the chart.
How many prompts and answers do I need for a reliable baseline?
More than most people expect. A single answer per prompt is noise. Our method collects about 100 real answers per prompt per market, which turns mention rate into a confidence band rather than a guess. The ladder we use is simple. 3 runs per prompt shows a pattern (the free tool), 5 to 10 runs is a DIY baseline, 20 runs is the minimum for a figure you would act on, and about 100 answers per prompt is what a full audit uses for a confidence band.
Find out what ChatGPT says about your brand before you spend anything.
Apply in about two minutes. We check your keywords, then run your real buyer prompts live on a 30-minute call. If we cannot create the result for your keywords, we tell you on the call, not after an invoice.
See if my keywords qualify →