Methodology
What we measure, and what keeps it honest.
A handful of rules govern every report we sell. The last is why you can trust the rest.
01
We measure answers, not rankings
A Saymetry run asks the AI engines the things real buyers in your market ask them, and records what comes back: every name recommended, every name mentioned, every source cited. How we design and target those question sets is our craft and stays ours; what you can always verify is the output, because your report shows every question asked, verbatim, next to how every engine treated it.
02
Seven engines, as buyers meet them
Every run covers ChatGPT, Claude, Gemini, Perplexity, Grok, DeepSeek and Mistral. Where an engine supports live web search, the run uses it, because that is how the deployed assistants actually answer buyers; where it does not, the answer is measured as the engine gives it, and the report says which was which.
03
Recommended, mentioned, cited: three different things
Being named is not being endorsed. We score every answer three ways: recommended (the engine steers the buyer to you), mentioned (you appear at all), and cited (your own pages are used as a source). The gaps between the three are usually where the actionable work hides.
04
Counts before percentages
Numbers are reported as plain counts first, not dressed up as percentages, because counts are checkable and percentages hide sample sizes. Where a run cannot classify enough answers to be trustworthy, it is held and re-run, never padded.
05
More answers is not more facts
Answers to the same question are related to each other. If a question is one your market simply does not associate you with, every engine tends to miss you on it, and those misses are not independent pieces of evidence, they are closer to one. Statisticians have measured this since the 1960s and it has a name, the design effect. We measure how clustered your own scan’s answers are and widen the confidence range to account for it, so the number we publish is honest about how much it actually knows. We have not found a competitor who publishes an interval at all.
06
What we will not tell you
A per-engine number needs far more questions than a total does, because each question is asked once per engine. At a real scan size a single engine’s rate would carry a range too wide to act on, so we never print a per-engine rate. We print the count and the number of questions behind it, and let you see the sample instead of dressing it up as a percentage. We would rather give you the few things we can stand behind than thirty we cannot.
07
How we decide something actually changed
Tracking asks the same fixed panel of questions every week, deliberately. Keeping the set fixed is what lets us tell a real movement from noise, because the same question compared against itself cancels out most of the randomness, change the questions and you are measuring the questions. We only call something a change when it clears the range; a movement of two or three points week to week is almost always the engines being engines, and we say so rather than drawing you a chart of it. We do not measure daily. Running the same questions every morning does not make the number more precise, it only gives you seven more chances a week to see a change that is not there.
08
The conflict firewall
Results are never adjusted for any commercial relationship, and nobody can pay to change a number. We sell the measurement and, separately, the work of improving your inputs; the measurement itself is not for sale, which is exactly why it stays worth paying for.
What the numbers mean
Every number, defined plainly.
You should know exactly what each figure in your report counts. As far as we can tell, no other AI visibility tool publishes a confidence range at all; we do, and here is what each number means.
- Recommendation rate
- Answers in which an engine names you as a pick, divided by answers we successfully analysed. Counted per engine and pooled. We report the count alongside the percentage, always.
- Share of recommendations
- Your recommendations divided by every recommendation counted across the whole field, times 100. Plain arithmetic on the numbers already in your report. Nothing is weighted, scored or modelled.
- Confidence interval
- Every rate carries a 95% confidence range. Answers are not independent, since the same question is put to every engine, so we account for how clustered your results are and widen the range rather than report a falsely narrow one. We would rather publish a wide honest number than a narrow flattering one, and we publish a range at all, which competitors do not.
- Agreement
- A stability check on how consistent this market’s answers are from one reading to the next. It tells you how settled the market is, and never touches your counts.
- Cited as a source
- Answers where the engine used one of your own pages while composing its reply, whoever it ultimately recommended. Being cited and being recommended are different things and we never merge them.
Limits
What can this measurement not tell you?
Answers are not deterministic.
Ask an engine the same question twice and the list of names can change. Published research finds fewer than one in a hundred repeated prompts returns an identical brand list. This is why we report a rate with an interval and refuse to report a ranking position: a position in an AI answer is not a stable quantity.
We query vendor APIs, logged out.
Every answer in your report was obtained through the official API as an unpersonalized request. A signed-in person with chat history can be answered differently. Your report states this on every run.
Engines disagree, and we never blend them.
Research finds the large majority of cited sources appear on exactly one engine. A single blended visibility score would hide that, so we report each engine separately.
Google AI Overviews and AI Mode are not in this scan.
They need search-results capture rather than an API, and we will say so rather than quietly imply coverage we do not have.
One run is a snapshot.
Trends are the signal. Any single figure, ours included, should be read with its interval, not as a precise measurement.
Why we publish this: a vendor once ran a controlled test on their own site where the pages they changed gained 56% more citations, and the untouched control pages gained 64%. Any AI visibility result without a control is indistinguishable from doing nothing. That applies to our numbers too, which is why yours arrive with an interval attached.