
Oct 5, 2026
What your ChatGPT visibility score really measures
By Amine Aziz Alaoui, LLM Research Lead, GetMint Research
This article summarises a research paper by Amine Aziz Alaoui, PhD in applied mathematics and LLM Research Lead at GetMint Research: A Measurement Model for Brand Visibility in ChatGPT Answers.
A growing part of product discovery now happens inside AI assistants like ChatGPT. Visibility there is becoming a priority for every brand. To improve it, you first need to measure it. Most tools give you a daily number without a reliable margin. Read that way, the number can mean very little. The risk is to spend time and money measuring and optimising a figure that says nothing about your brand. This article shows why, and how to avoid it.
Full paper: A Measurement Model for Brand Visibility in ChatGPT Answers (preprint v1.3, 5 October 2026).
Your AI visibility score is the share of ChatGPT answers to a question that name your brand. It jumps every day, even when nothing has changed.
ChatGPT has a mood of the day, for each question. We call it the weather of the day. Ask the same question 100 times today and you get one score. Ask it 100 times again tomorrow and you get a different score. Your brand did nothing different. ChatGPT simply behaves differently that day: it picks other sources and writes its answers another way. On one question, this mood moves the score by about 14 points from one day to the next. It fades within two to three days.
Your daily score mixes four things. Your real visibility, the weather of the day, ChatGPT's updates, and chance. Only your real visibility is worth tracking.
Repeating the question on the same day does not remove the mood. All the answers to one question on one day share largely the same mood. So the score of one question read on one day can be off by about ±28 points, even with 1,000 answers.
Precision comes from more questions and more days. Each question and each day has its own mood, so the moods cancel out when you combine them. About 20 questions close in meaning, all targeting the same need, each asked once a day, read over 14 days: the score is reliable to about ±7 to ±8 points outside ChatGPT update days.
ChatGPT changes without warning. From May to July 2026, we identified 7 days when the brands we track moved together. 6 had no public announcement nearby.
What it changes for you: decide on groups of about 20 questions that stay close in meaning and express the same need or search intent, read two weeks of data, ignore moves inside the margin, and check for a ChatGPT change before reacting to a drop.
Why your score jumps every day
Your AI visibility score was 46% yesterday. Today it reads 41%. Nobody published anything. No competitor launched. If you track your brand in ChatGPT, you have lived this.
A visibility score is the share of ChatGPT answers to a question that name your brand. ChatGPT never answers the same way twice. Ask the same question twice and your brand may appear in one answer and not in the other. So your score is an estimate, like a poll. A poll says "30%, give or take 3 points". Your score needs that margin too. Few tools show it: a 2026 comparison of nine monitoring tools found two that publish one.

This matters for your decisions. Take a topic of 20 questions, each asked 30 times on the same day. Most tools would show a margin of ±4 points, because they only count chance. The real margin is about ±7, because the weather of the day comes on top. A move that looks real on their dashboard can be pure weather. Asking each question more than 30 times that day barely helps: with 100 answers per question, the margin is still about ±7. You pay for more than three times as many answers and gain almost nothing. And 6 of the 7 ChatGPT changes we identified came with no public announcement. Without the right margin and a way to spot these changes, you react to moves that have nothing to do with your work.
GetMint Research studied eight months of daily monitoring of ChatGPT's web interface, from January to September 2026. That is more than a million answers to tens of thousands of questions. Here is what we found, without the maths.
What we found: four things move your score
Think of weather and climate. The weather changes every day. The climate is what you live in, and it shifts gradually. ChatGPT's mood of the day is the weather. Your real visibility is the climate. Your score has both, plus two other sources of movement.
What moves | What it is | Typical size, one question | What helps |
|---|---|---|---|
Your real visibility (the climate) | How often ChatGPT names you once the noise is removed. It moves on its own and with your actions. | About 16 points over two weeks | Nothing: this is what you want to track |
The weather of the day | What ChatGPT does that day for that question: which sources it retrieves, how it writes | About 14 points from one day to the next. Gone within 2 to 3 days | More days and more questions |
ChatGPT updates | Days when ChatGPT changes something for all brands at once | About one every 13 days from May to July 2026. One episode measured at 15 to 18 points | Detecting and flagging these days |
Chance | Each answer is a random draw | Shrinks as answers add up | More answers |
Points are percentage points of visibility. Sizes are measured on questions where the brand is neither almost always nor almost never named.
The surprise is the size of the weather. Take one question read on two days a few weeks apart. The two readings typically differ by 25 to 33 points. About 20 of those points are the weather of each day. A 5-point drop from one day to the next tells you nothing about your brand.

Most tools compute their margin from chance alone. That margin looks precise and is too narrow, because it ignores the weather.
What it means for how you measure
1. Precision comes from questions and days
Repeating a question on the same day reduces chance, up to a point. It never removes the weather of that day. One question read on one day stays uncertain by about ±28 points, whatever the number of answers. In plain words: the real value can sit anywhere from 28 points below the reading to 28 points above it. A single question never gets much below ±30 points, even over many days.

So group about 20 questions into a topic. A topic is a set of questions that stay close in meaning and express the same need or search intent, for example "best running shoes for beginners", "which running shoe should I buy first", "good shoes to start running". Each question has its own weather. Outside ChatGPT update days, their daily moves are almost unrelated, so they cancel out.

Then read several days. For a topic of 20 questions, each asked once a day:

Two weeks is the sweet spot. Your real visibility keeps moving, so older days describe a visibility you no longer have. In our simulations, a 30-day average shown with a chance-only margin contained the real visibility on only 58% of days without a ChatGPT update.
Repeats still help a topic, up to a point. 20 questions with 30 answers each on one day (600 answers) also give about ±7. Beyond 30 answers per question, extra answers add nothing. Fourteen days at one answer per question per day reach about the same precision with 280 answers. The paper combines both: many answers in one day to set a starting point fast, then one answer a day.
2. ChatGPT updates must be flagged
We built a daily check, the Sentinel, that spots days when tracked brands move together. Between May and July 2026 it identified 7 such days. 6 had no public announcement within two days. One of them may come from a change in our own measurement tool. On 8 August, ChatGPT silently switched to a new model version. That day showed the largest common move since late July. The public launch of GPT-5.6 on 9 July showed no detectable effect on brand mentions. Announcements and real impact are two different things.

If these days stay hidden inside the average, the margin of a 20-question topic over 14 days can grow from about ±7 to about ±13 points. Once a day is flagged, the score is recalculated starting from that day.
3. Judging your actions takes two weeks and a comparison
Once the weather and the updates are set aside, what remains is your real visibility. That is where your actions show. Two cautions apply. A real rise shows up late: in our simulations, every method lagged the real visibility by about a week. And a rise alone proves nothing, because your real visibility also moves on its own. The paper proposes to compare your questions with similar questions from other brands over the same period, outside flagged days, and to check that your own pages appear among the sources ChatGPT uses. Their application to client cases is a separate study, not reported there.

Why some scores look stable when they are not
If each question is asked only once a day, the standard way to calculate how much a score moves makes it look stable for about two weeks. Our own earlier analysis made this mistake: days with several answers per question show the score moving from the very first day.
Six rules to read your score
Decide on topics: about 20 questions close in meaning, targeting the same need. A single question is too noisy to decide anything.
Read the last 14 days. Yesterday's number is mostly weather.
Ignore any change smaller than the margin.
Before reacting to a drop, check whether ChatGPT changed that day.
Give an action about two weeks before judging it, and compare it with other brands over the same period.
Stop adding answers on the same day beyond about 30 per question. Add days and questions instead.
Read the paper
A Measurement Model for Brand Visibility in ChatGPT Answers, Amine Aziz Alaoui, GetMint Research, preprint v1.3, 5 October 2026. DOI: 10.5281/zenodo.23156371.
The aggregate tables and the analysis code are available from the author. A public study package will come with the archived version. Raw answers stay private. The study has some limits, which you can read in the full paper. Read the paper, challenge it, and tell us where we are wrong.
START FREE. NO CREDIT CARD REQUIRED.
Your competitors are already visible in AI. Are you?
See what ChatGPT, Gemini and Perplexity say about your brand. Set up in 5 minutes.

