This is the question a careful buyer asks before paying for any AI visibility tool, and most vendors dodge it. Here is the straight answer. AI answers are probabilistic, so any single measurement is noisy. Accuracy does not come from a cleaner algorithm. It comes from sampling the same questions more than once, covering many engines, and reporting a confidence range instead of one hard number. A tool that hands you a single tidy score with no uncertainty attached is overclaiming, and you should treat that as a warning sign.
The rest of this page is the long version. It walks through where the noise comes from, how honest measurement keeps each source of it in check, and how to read a report critically enough that you could verify the numbers yourself before you trust them.
The error range, worked through on a real number
Here is what a plus-or-minus actually means in practice. Suppose a scan asks 60 buyer questions across 7 engines and gets about 350 usable answers, and your business is recommended in 71 of them. The raw rate is 20.3%. A 95% Wilson interval on that count runs from roughly 16% to 25%, and because several answers to the same question are correlated rather than independent, an honest tool widens that band further using a design effect measured from the scan's own answers. So the truthful report is not "20.3%" but "around 20%, and we would not be surprised by anything from the mid-teens to the mid-twenties." Next week the raw number reads 22.1%. Is that progress? No: it sits comfortably inside the band, so an honest report calls it noise. Only a move that clears the interval is a real change. This is the arithmetic behind every number Saymetry publishes, formula included, on the methodology page, and it is the difference between measurement and a mood ring.
Why one measurement is not enough
The engines do not return the same thing twice. Ask ChatGPT or Gemini the same buyer question a minute apart and the list of names, and the order they come in, can move. This is not a bug in the tracking. It is how the models work. Published research on repeated prompting finds that fewer than one in a hundred repeated prompts returns an identical list of brands. So the first thing to understand is that AI visibility is a probability, not a rank. A brand that shows up in four of ten runs has a 40 percent mention rate. Any tool reporting a single daily "position" for that brand is reporting noise dressed up as a fact.
The five sources of variance
The movement in an AI answer is not one thing. It has at least five distinct sources, and a measurement is only as honest as its treatment of all of them. Here is each one with a concrete example.
- Repeat variance. The same question, asked again, returns a different set of names. Ask "best project management tool for a small agency" twice and you might get Asana, Trello, and ClickUp the first time, then Asana, Notion, and Monday the second. Nothing changed but the roll of the dice. This is the largest source, and it is why a single pull tells you almost nothing.
- Engine variance. ChatGPT, Claude, Gemini, Perplexity, Grok, DeepSeek, and Mistral do not agree with each other. A brand that ChatGPT names in half its answers can be absent from Gemini entirely. Research on the sources these engines cite finds the large majority appear on exactly one engine, not shared across them. Visibility on one says very little about the others.
- Time variance. Models get updated, indexes refresh, and the web the engine reads keeps moving. A brand that a model surfaced strongly in July can slip in August because the underlying model was retrained or a competitor picked up a run of new mentions. The answer this month is not guaranteed next month.
- Personalization variance. A signed-in person with months of chat history and stated preferences can be answered differently from a cold, logged-out request. Someone who has been discussing budget tools all week may get a budget-skewed shortlist that a fresh session never sees.
- Prompt-wording variance. The words you use change the answer. "Best CRM", "top CRM for startups", and "affordable CRM for a two-person team" are three different questions that return three different shortlists. A tool that quietly reworks its prompts between runs is measuring the wording, not your visibility.
How good measurement controls each one
None of that variance means tracking is hopeless. It means the method has to account for it out in the open, rather than hiding it behind a single number. Each source of noise has a specific control, and an honest tool names the control it uses.
- Repeat variance is controlled by repeated sampling. A share of every scan is asked a second time on every engine. That second pass is how you measure how stable the market is, and how much of any change is just randomness rather than a real shift.
- Engine variance is controlled by keeping engines separate. Because the engines disagree, a single merged score would hide the disagreement. Each engine is reported on its own, so you can see where you are strong and where you are missing entirely instead of averaging the two into a meaningless middle.
- Time variance is controlled by a fixed panel and trend reading. The same questions are reused every run, so a movement compared against itself cancels out most of the randomness. One run is a snapshot; the trend across runs is the signal. Change the questions and you are measuring the questions.
- Personalization variance is controlled by logged-out requests. Every answer is pulled through the official API as a cold, unpersonalized request, so results are not skewed by one account's history. The report states this on every run, because a signed-in person can and does see something different.
- Prompt-wording variance is controlled by holding the wording fixed. The panel is written once and reused verbatim, and every question is shown in your report. If the words never change between runs, a change in the result is a change in your visibility, not a change in the prompt.
On top of those controls sits the number itself. It arrives with a confidence range around it, because answers to the same question are related to each other and cannot be treated as independent coin flips. One question is asked of every engine, and a question your market clearly associates with you tends to be recommended across many of them while an irrelevant one is recommended nowhere. Treat those as independent and the interval you print is too narrow to trust. Our method measures how clustered your own scan is and widens the range to match, which is why the hundreds of answers in a scan are treated as a much smaller pool of independent ones. The formula is published on our methodology page. We would rather show a wide honest number than a narrow flattering one.
Noise source versus control, at a glance
| Source of noise | What it looks like | How it is controlled |
|---|---|---|
| Repeat variance | Same question, different names each time | Ask a share of questions twice, measure agreement |
| Engine variance | Strong on ChatGPT, absent on Gemini | Report every engine separately, never blended |
| Time variance | A model update reshuffles the answer | Fixed panel across runs, read the trend not the snapshot |
| Personalization variance | Signed-in history skews the shortlist | Logged-out API requests, stated on every run |
| Prompt-wording variance | Reworded prompt returns a new list | Wording held fixed and shown verbatim |
What "accurate" should mean here
Accurate does not mean one precise number. Nothing about a probabilistic system supports that. Accurate here means directionally reliable and repeatable, with the method shown so you can verify it yourself. Two things make that possible. Every question we ask is in your report, verbatim, next to how each engine answered it. And the method is published, formulas included, so you are not asked to take the number on faith. It also means counts before percentages. "Recommended in 17 of the questions we asked" is checkable and carries its own sample size. A bare 28 percent with no denominator hides how much evidence sits underneath it.
How to read a report critically
You do not need a statistics background to pressure-test one of these reports. You need to ask five questions, and you can answer most of them in ten minutes.
- Are the raw questions shown? Read them. If a tool hides its prompts as a "secret sauce", there is nothing for you to check, and the number could be built on anything.
- Is there a range on the score? A point estimate with no interval is a marketing claim, not a measurement. The interval is where the honesty lives.
- Are the engines kept separate? A single blended visibility score averages away the exact disagreement you need to see. Look for a per-engine breakdown, or at least a count per engine.
- Was anything asked more than once? If every question is asked a single time, the report cannot tell you how much of your movement is noise, because it never measured the noise.
- Does it state what it cannot measure? A serious tool tells you up front that it queries logged out, that Google AI Overviews are not in the scan, and that a guaranteed recommendation is not something anyone can sell. Silence on the limits is itself a signal.
Then verify it yourself. Take three of the questions from the report, open ChatGPT, Gemini, and Perplexity in a logged-out window, and run them a few times each. You will watch the names move between runs. That is the variance the report is supposed to be handling, and now you can see whether it did.
Honest measurement versus overclaiming
| Sign it is honest | Sign it is overclaiming |
|---|---|
| Reports a range or interval on the score | One clean number, no uncertainty shown |
| Shows you every question it asked | Hides the prompts as a "secret sauce" |
| Covers several engines, kept separate | One engine, or a single blended score |
| Samples repeats to measure stability | Asks once and calls it the answer |
| Reports counts with the sample size | Prints a percentage with no denominator |
| Says plainly what it cannot measure | Promises a guaranteed recommendation |
Be candid about the limits
No one can guarantee that an engine recommends you. Anyone who promises that is selling something they do not control. The models are not deterministic, and their makers change them without notice. What is measurable, and worth paying for, is different and more useful: your presence in the answers, your share of voice against the rest of the field, and the sources the engines lean on when they answer. Those are checkable, repeatable, and directly tied to work you can actually do. A promised ranking is not.
There are things this kind of measurement cannot tell you, and a serious tool states them before you find them. A single run is a snapshot, and trends are the real signal. Logged-out API answers can differ from what a signed-in person sees. Google AI Overviews and AI Mode need search-results capture rather than an API, so they are not in this scan, and we say so rather than imply coverage we do not have. And any single figure, ours included, should be read with its interval, not as a hard measurement. We would rather publish a wide honest number than a narrow flattering one.
Frequently asked
Why did my score change when nothing changed on my end?
Because the engines are probabilistic. Ask the same question twice and the list of names can come back different, with no change to your site at all. A week-to-week move of a few points is almost always the engines being engines, not a real shift. That is why a good tool reports a range and only calls a change when the number clears that range.
How many times do you ask each question?
A share of every scan is asked a second time on every engine, so we can measure how stable the market is and how much of any movement is noise. A fixed panel of questions is reused week to week so a change is a real change, not a reworded prompt.
Can you guarantee ChatGPT will recommend me?
No, and anyone who promises that is lying. No one controls what an engine returns. What is measurable is your presence, your share of the recommendations, and the sources behind the answers. We report those with a confidence range, and we show every question we asked so you can check the work.
So what does "accurate" actually mean here?
Directionally reliable and repeatable, with the method shown so you can verify it. It does not mean one hard number carried to two decimals. It means the same panel, sampled more than once, across many engines, reported as a range you can trust.
Why report a raw count instead of a clean percentage?
Because a count is checkable and a bare percentage hides the sample behind it. "Recommended in 17 of the questions we asked" tells you both the result and how much evidence sits under it, so you can judge for yourself whether a move is worth acting on. A tool that shows you 28 percent with no denominator is asking you to trust a number you cannot audit.
How do I verify the tracking myself?
Open the report and read the raw questions, verbatim, next to how each engine answered them. Pick a few and run them yourself in ChatGPT, Gemini, and Perplexity while logged out. You will see the same variance we do, and you will see whether the report accounts for it with a range or hides it behind one tidy figure. If the questions are not shown, there is nothing to check, and you should treat that as the answer.