A visibility score is an important view into how AI assistants surface your brand. They can vary a lot, however.
Imagine that your visibility score was 46% yesterday. This morning it’s 41%. Nobody shipped anything, no competitor launched, and your business hasn’t changed. So you spend the day digging through the web for answers but find nothing. The next morning, the score is 45%, as if nothing had happened.
That swing has an explanation, and it usually isn’t you. I analyzed seven months of visibility data across 120 million chats to find out which moves are real and which are noise. Telling these movements apart contains important lessons for ensuring this central metric is trustworthy and useful for guiding decision-making.
TL;DR
Choose more prompts rather than more runs to improve your visibility score’s accuracy. The prompts you track are only a sample of all the questions asked by your customers every day. Asking many questions once tells you more than asking one question over and over.
Time is your friend for detecting score changes. Even when tracking only 50 prompts, you can detect real visibility changes. Time makes the difference. The longer you wait the more confidently you can know visibility has moved.
Differentiate between branded and non-branded prompts for best results. The visibility of branded prompts rarely changes. Non-branded prompts are the place where optimizations have the largest impact.
Every prompt is a biased coin; or why visibility appears to vary
Imagine a coin that lands on average 8 times out of 10 on heads. This would be an unfair, biased coin.
When you flip this coin once, you’ll see either heads or tails. But you’ll get no information about how biased the coin is. Only by flipping it a lot will you start to see that heads comes up about 80% of the time. The more you flip it, the more confident you’ll be, and you’ll probably start assuming the coin’s weight is off.
AI engines answer your tracked prompts exactly like this – with strange weight. The wording of a prompt influences the coin, setting a fixed rate at which the engine mentions your brand when asked multiple times. Every run is one flip. Heads, you're mentioned. Tails, you're not. And the results aren’t 50/50.
Branded prompts are decided, non-branded prompts are not
You can separate prompts into two types, since they behave nothing alike. Branded prompts name a company: “Does VW make a good electric car for families?” Non-branded prompts ask open questions: "What is the best electric car for a family of four?"
Branded prompts are decided. Ask about a company, and the engine mentions it run after run. Non-branded prompts are uncertain, and the engine will search the web and, occasionally, its training data without having a company already suggested.
Will your brand get mentioned in response to a non-branded prompt? The answer is neither a clear yes nor a clear no. Not every run of the prompt will mention the same brands. If you check a non-branded prompt a few times, the results may even look random. While you may think that’s an error, it isn’t. Rather, the seemingly random results are due to the underlying visibility rate being closer to even. In those cases where visibility is close to 50/50, you need more samples before you can tell where your brand’s true visibility stands.
How many runs do you need until you know a prompt’s visibility level?
The number of runs you need depends on how close the prompt's true visibility is to 50%. Say you let one prompt run once a day for 14 days. If your brand shows up in all 14 responses, you can already be confident the true visibility is at least 80%. But if it shows up in 6 of the 14, that tells you almost nothing. The true visibility could be anywhere from about 20% all the way up to 70%.
The closer your brand’s true visibility sits to 50% for a given prompt, the more runs it takes to know if a visibility score is accurate or just chance. The more contested a prompt is among brands, the less you know whether you improved, fell behind, or fell into the result by chance.
There is a formula to calculate how far off the visibility score can be, called the Wald confidence interval:
how far off visibility can be = 1.96 × √(( visibility × (1 − visibility)) ÷ number of runs)
Where “How far off visibility can be” is expressed as percentage points (pp). Importantly, this is only an approximation, and the confidence interval doesn’t hold when visibility is close to 0 or 1. (In these cases with visibility close to the minimum or maximum, you can use the rule of three: how far off you can be = 3 ÷ number of runs)
If we use the Wald confidence interval, the math tells us that to get an estimate of your brand’s visibility score that is two times more precise, you need four times the runs. Pinning a prompt with a true visibility of 50% down to ±10 pp takes about 100 runs. Down to ±5 pp, almost 400.
How many runs until you know a prompt’s visibility has changed?
Getting more precise estimates of confidence intervals assumes that the true visibility score remains the same. But of course a prompt’s visibility score doesn’t have to stay fixed. You improve your website, create content, and reach out to customers for feedback, all to shift visibility in your favor.
That effort takes time to show up in the numbers.
Knowing the visibility level already takes a lot of runs, and confirming the changes takes roughly double the runs again.
If you’re only tracking one prompt at one run a day, it would take you three months to a year of tracking to be certain about your brand’s visibility score, and about six months to two years to be certain about how much it changed.
Tracking only one prompt would therefore be highly inefficient. Nobody runs only one prompt hundreds of times, and neither do we. A single prompt, read alone, would keep your brand’s visibility level and any change in it mostly hidden for a long time.
Your prompt portfolio is only a sample of the market
We don’t just look at one prompt to see how often a brand is mentioned by AI assistants. We look at a portfolio of multiple prompts that we deem relevant.
Better coverage introduces more uncertainty, though. Every day, millions or even billions of questions are asked across the globe. Tens of thousands are probably relevant for your brand, but your portfolio only tracks a very small fraction of them. This means the visibility score for the prompts in your portfolio is only an approximation of the true visibility across all possible prompts run per day.
If every prompt’s actual visibility were equal, running one prompt 100 times would tell you the same thing as running 100 different prompts once. But because each prompt is different, your brand’s visibility will vary between them. The average visibility across your portfolio therefore depends on which prompts you choose to track. This is called sample variance. Crucially, it depends only on how many prompts you track, not on how often you run them
Portfolio hygiene also affects visibility score accuracy
The way you manage your prompt portfolio can influence how accurate your brand’s visibility score is. If you aren’t grouping your prompts, it can affect how well your visibility score actually answers the question of whether your brand appears when customers prompt.
The most important separation you can introduce is one between prompts that mention your brands and the ones that don’t, since they behave differently. As mentioned, branded prompts impose a path dependency that makes the chances of your brand appearing in the results extremely high. Grouping these separately from non-branded prompts — which are genuinely contested — gives you a more accurate picture of your brand’s appearance in AI assistant results.
Is it better to run fewer prompts more often, or more prompts less often?
To answer the question of how many prompts your portfolio should track, I want to look at two scenarios where a brand is trying to ascertain its visibility score. In both cases you have 1,500 runs to allocate per month for one AI assistant. 20% of your prompts are branded and 80% are not.
Scenario 1: You run 50 prompts once per day (50*30*1= 1,500).
Scenario 2: You run 10 prompts five times per day (10*30*5=1,500).
For the calculation of the confidence intervals, I used numbers from our database across ~120 million chats between February and August 2026 on every major AI assistant provider (ChatGPT, AI Mode, AI Overviews, Gemini, Perplexity, Copilot, Grok, and Claude).
The variance for branded prompts is ~0.0225 and for non-branded prompts is roughly 3 times as much at ~0.07. Our database shows 40% of non-branded prompts are contested (visibility between 10% and 90%), versus just 10% of branded prompts.
How precisely can you know your brand’s visibility score in each scenario?
In scenario 1, where you run 50 prompts daily, the average visibility level based on a given day could be ±10.5 pp at a 95% confidence level. But at the end of the month, when you consider the monthly average, you’ll find the uncertainty has reduced to ±6.1 pp as you gather all 1,500 samples.
Scenario 2 allocates the same number of samples per day, but the daily visibility can vary +- 15.8 pp — roughly 50% higher than scenario 1. At the end of the month, the uncertainty would be +-13.2 pp, which is more than double the uncertainty for scenario 1. The reason for this is the sample variance. If you only track a tiny fraction of all questions asked per day, how many questions you ask is more relevant than how often you ask them.
In both scenarios, to be roughly twice as precise you have to quadruple the number of questions you ask. Increasing the amount of runs has nearly no impact.
This shows two things. First, when tracking a smaller number of prompts, the visibility score amounts to little more than an estimate, highly affected by chance. Swings should be expected, especially when looking at daily timeframes. Second, to identify the visibility score more precisely, your only option is to track more prompts. Increasing run frequency won’t help.
When do you know that your visibility has changed?
Unlike determining the certainty of your brand’s visibility level, detecting change depends only on how many runs you have, not on how many prompts you track. That’s because you’re comparing the same prompt against itself over time.
In both scenarios you can be equally certain after the same number of runs. The more samples you collect, the more precisely you can detect changes in the visibility score.
Daily changes below 12.4 pp could be due to chance, since you only collected 50 runs. However, by looking at the weekly visibility score with 350 runs, or the monthly score with 1,500 runs, you can be certain that changes above 4.7 pp (weekly) and 2.3 pp (monthly) represent real shifts.
This shows that you can be very precise even with a small number of prompts, as long as you keep tracking the same prompts.
So let’s consider the imagined situation that opened this report, where your brand’s visibility score jumped from 46% to 41% to 45% in 48 hours. Under both scenario 1 and scenario 2, the likelihood of this jump is 11%. But considering such a jump on a weekly timeframe, where your prompt tracker has completed more runs, the likelihood drops to 0.56%. On the monthly timeframe, the chances of such bounces are roughly one in 1.4 million (0.00007%).
Putting it into practice: five rules for your dashboard
This research points to a few useful habits that can improve the accuracy and actionability of your brand’s visibility score:
Step 1. Group your prompts. Branded and non-branded prompts behave differently, so always tag them, and don’t compare them in one number. In Peec AI, this is done automatically.
Step 2. Don’t look at single prompts for visibility. If you take one thing from this article, it should be this: Look at the portfolio or sub-groups, which you can filter on Peec AI using topics and tags. Not simple prompts. A single prompt can be very imprecise, which could lead to misleading interpretations when you get (un)lucky.
Step 3. Judge movement on the portfolio across time. A single day of data depends a lot on chance and can therefore be misleading. Looking at monthly data instead gives you a more reliable picture. In Peec AI, you can aggregate visibility trends across timeframes, AI models, topics, and tags, so you know your performance in each segment.
Step 4. Add prompts before adding runs. Tracking more prompts gets you reliable data faster than running the same prompts more often, and it also lets you filter more detailed views when looking at longer timeframes. The goal is to know what the market sees, not to perfectly understand a single question.
Step 5. Freeze the prompt set. Avoid constant changes to your prompt set. If you compare timeframes, make sure to only look at prompts that were tracked during the whole period.






