Deep article

How to Run a ChatGPT Visibility Audit Yourself (Step by Step)

A ChatGPT visibility audit answers one question: when buyers ask ChatGPT for a recommendation in your category, how often are you named, where in the list, and who beats you? You can run a credible small version in an afternoon with a prompt cluster, a spreadsheet and repeated runs. This guide gives the steps and the limits.

7 min readPublished October 8, 2026Published by Odys Global

You can find out whether ChatGPT recommends your brand in an afternoon, with nothing more than a list of buyer prompts, a spreadsheet and the discipline to run each prompt more than once. The result will not be precise, but it will be honest, and it will tell you which of your buyers’ questions you are winning, losing or absent from. Below is the exact procedure, followed by how to read the numbers and where a DIY audit stops being enough.

What a visibility audit measures and why one prompt is not an audit

A ChatGPT visibility audit measures four things across a set of buyer prompts: your mention rate (how often you are named), your position when named, whether a citation to your domain appears, and which competitors are named. Measured across a prompt cluster and many runs, those four numbers describe your real standing.

The reason you cannot shortcut this is that ChatGPT’s answers vary. The same prompt asked ten times produces different shortlists, and different phrasings of the same intent produce different shortlists again. Our own audit methodology collects around 100 real answers per prompt per market from genuine mobile connections, because that volume turns a guess into a confidence band. A DIY audit uses fewer runs and accepts wider error, which is fine for a first baseline as long as you read it that way.

Step 1: Choose the buying intents that matter

Do not audit everything. Pick two or three buying moments where a recommendation changes the outcome. For most companies that is the “best X for Y” question, the “X vs Z” comparison, and the constraint question (“X that does W”). Write the plain version of each.

Illustration: a payroll software company for restaurants might choose “best payroll software for restaurants,” “Gusto vs alternatives for restaurant groups” and “payroll software that handles tips and multiple locations.” Each is a different moment with a different cast of competitors.

Step 2: Build a prompt cluster of 20 to 40 variations

Expand each intent into the forms real buyers type. Vary six things: who is asking (owner, finance lead, operations manager), size or budget, a constraint (integration, compliance, price), location, comparison frame and question style. Five to eight variations per intent gives you 20 to 40 prompts in total.

Write them as full sentences, the way people talk to a chat box: “I run three restaurants in Texas with about 60 staff, what payroll software should I use?” rather than “restaurant payroll texas.” The method for building and pruning the list is in prompt clusters, and our ChatGPT prompt generator will produce a starter set of five for your category, location and buyer type.

Step 3: Set up the scoring sheet

One row per answer, not per prompt. Columns:

Column Values
Prompt ID and text P01 to P40
Run number 1 to 5 (or 10)
Date and time So you can see drift
Search triggered Yes if a sources panel or citations appear
Your brand mentioned Yes or no
Your position 1, 2, 3… or blank
Link to your domain Yes or no
Competitors named, in order Comma-separated
Notes Wrong description, caveat, hedging

Keep it this simple. The value is in consistency across runs and months, not in clever fields.

Step 4: Collect answers under consistent conditions

Conditions change answers, so fix them before you start. Use a logged-out session or an account with memory off, start a fresh chat for every run, and if you can, use a phone on a mobile connection rather than a desktop on office wifi. Answers differ by country, so if your market is the UK, collect from the UK.

Run each prompt at least five times, ten if you have the patience. Paste each answer into the sheet and classify it immediately while the context is fresh. Note whether ChatGPT searched the web, because searched and memory-based answers behave differently and reward different fixes. For 30 prompts at five runs this is 150 answers, roughly three hours of work.

Step 5: Score mention rate, position and share of voice

With the sheet filled, compute four figures:

  • Mention rate: answers naming you divided by total answers, overall and per prompt.
  • Average position when mentioned, and top-three rate (answers where you are in position 1 to 3 divided by total answers).
  • Citation rate: answers with a link to your domain divided by total answers.
  • Share of voice: each competitor’s mention rate alongside yours.

Illustration: suppose 150 answers name you in 48, place you in the top three in 27, and link your domain in 6. Your mention rate is 32%, your top-three rate is 18%, your citation rate is 4%. If the leading competitor is named in 105 answers (70%), you now know the size of the gap. Why position deserves its own number is covered in why position 1 to 3 in a ChatGPT answer is the whole game.

Step 6: Run the direct entity test

Separately from the buyer prompts, ask ChatGPT about you by name: “What is [Brand] and what does it do?”, “Who is [Brand] for?”, “Who are [Brand]’s main competitors?”, “Where is [Brand] based?” Run each three times. If the model misdescribes you, confuses you with another company or lists competitors from the wrong category, your entity is weak, and that usually explains a low mention rate on its own. The fix sequence is in brand entity consistency for LLMs.

Step 7: Diagnose the gaps prompt by prompt

Sort the prompts into three groups: top three, mentioned but lower, and absent. Then look at what the groups have in common. Do you lose every prompt with a budget constraint? Every prompt phrased in buyer language rather than your positioning? Every prompt with a location? Which competitor wins the prompts you lose?

Illustration: suppose you are in the top three for 8 prompts, mentioned lower for 9 and absent from 13, and the absent group is dominated by prompts containing “affordable” or “for small teams,” with the same two competitors winning them. That is a precise brief: the market does not describe you as the small-team option, and the sources ChatGPT reads say those two competitors are. The seven common causes of absence are mapped to symptoms in why ChatGPT doesn’t mention your brand.

Step 8: Record the baseline and schedule the re-audit

Freeze the prompt list in writing. Save the sheet as “Baseline, October 2026.” Put a calendar entry 30 days out to repeat the identical process: same prompts, same conditions, same number of runs. A re-audit is only meaningful if nothing about the measurement changed except the date.

Resist the urge to tweak prompts between audits. If buyer language genuinely shifts, add new prompts as a separate set and keep the original set running for comparison.

Reading a DIY result honestly

Five runs per prompt gives a wide margin. A prompt-level mention rate of 40% on five runs could easily be 20% or 60% on the next five. Trust cluster-level figures more than prompt-level ones, treat month-over-month changes under about 20 points as noise, and look for direction across several prompts rather than movement on one. The statistical reasoning, and why around 100 answers per prompt is where numbers become decision-grade, is set out in measuring ChatGPT visibility by mention rate and position.

Also be aware of what a laptop in one city cannot see: answers in your target country if you are elsewhere, differences between free and paid tiers, and differences between mobile and desktop. Those are the reasons a professional baseline exists.

When to move from DIY to a full audit

A DIY audit is the right first step and may be all a small business needs. Move to a full audit when you need numbers you can report to a board, when you operate in more than one country, when you are about to invest in a visibility program and need a defensible before-and-after, or when your DIY result sits in the ambiguous middle (20 to 50% mention rate) where five runs cannot tell improvement from noise. The trade-offs between tracking software and a done-for-you audit are compared in AI visibility tools vs a done-for-you audit.

Our baseline audit collects around 100 answers per prompt per market from real mobile connections, reads and classifies every answer, and reports at the prompt level so you can reproduce any result yourself in ChatGPT. If you want that instead of the spreadsheet, find out what ChatGPT tells buyers about you.

What to do next

  • Write your three most commercial buying questions and expand them to 20 to 40 prompts today.
  • Run each at least five times logged out, score every answer, and compute mention rate, top-three rate and share of voice.
  • Save the baseline, diary the 30-day re-audit, and start with the entity fix if the direct test showed inaccuracies.

Frequently asked questions

How do I check if ChatGPT recommends my brand?

Ask it the way a buyer would, not by name. Type 20 to 40 buyer-style prompts for your category ("best X for a small business in 2026"), run each several times in fresh chats, and record whether your brand is named and in what position. One prompt run once is not a check; answers vary between runs, so the honest result is a rate across many answers.

How many times should I run each prompt?

Use the ladder. 3 runs per prompt shows a pattern (the free tool), 5 to 10 runs is a DIY baseline, 20 runs is the minimum for a figure you would act on, and about 100 answers per prompt is what a full audit uses for a confidence band. With five runs, treat differences smaller than about 20 percentage points as noise.

Should I run the audit logged in or logged out?

Logged out, in a fresh chat, on a mobile connection if you can. Logged-in accounts carry memory and chat history that personalize answers toward brands you have discussed. Buyers who have never heard of you have none of that context. If you must use an account, turn off memory and start a new chat for every run.

What tools do I need for a DIY ChatGPT audit?

A ChatGPT account or logged-out session, a spreadsheet and two to four hours. No paid software is required for a first baseline. Tracking tools become useful when you want daily sampling across many prompts and markets, but they do not change the method, which is to collect many real answers and classify each one.

What are the limits of a do-it-yourself audit?

Sample size, location and consistency. Five runs per prompt from one laptop in one city tells you roughly where you stand, but not whether a 10-point change next month is real. Answers also differ by country, device and account state. A DIY audit is the right first step; a professional baseline with around 100 answers per prompt from real mobile connections in your target market is how you get numbers you can act on.

Your next move

Find out what ChatGPT says about your brand before you spend anything.

Apply in about two minutes. We check your keywords, then run your real buyer prompts live on a 30-minute call. If we cannot create the result for your keywords, we tell you on the call, not after an invoice.

See if my keywords qualify →
2-minute application30-minute live audit callWe decline keywords we cannot move

Playbooks by industry

Check my keywords →