Citeify
← All news
News
5 August 2026

Ask ChatGPT the same question twice: why AI visibility scores wobble, and how to measure honestly

AI engines give different answers to the same question on different runs — so a single screenshot proves nothing. How Citeify measures AI visibility with repeated sampling, majority voting and a frozen question set.

Run the same buyer question through ChatGPT twice and you will often get two different answers, citing different sources. Recent research measuring day-to-day overlap in generative engines' cited sources finds that typically only around a third of sources persist between runs — the rest churn. Academic work on GEO measurement now recommends multiple repetitions per question before treating any visibility result as real.

This has an awkward implication for the screenshot economy: that image of ChatGPT recommending a brand — the one in the agency pitch deck — is one sample of a distribution. It proves roughly as much as a single coin flip proves a coin is biased.

What honest measurement looks like

When Citeify measures AI visibility, three rules apply:

1. Repeated sampling, majority vote. Every question is put to every engine multiple times. A brand only counts as present or cited if it appears in the majority of samples. Marginal, lucky appearances get filtered out — which means the score is lower than a cherry-picked screenshot, and much more likely to be true.

2. A frozen question set. The first audit of a domain generates buyer-intent questions from what the site actually does. Every subsequent audit of that domain re-asks the same questions. Without this, a score change between two audits could mean your visibility moved — or just that different questions were asked. With it, a delta means what it appears to mean.

3. Per-engine reporting. Cross-platform citation studies find that only a small minority of domains are cited by more than one AI platform for the same query — each engine has its own sources and its own habits. One blended number hides that; we report each engine separately, then aggregate transparently.

The trade we're making

Measured this way, scores are less flattering. Our own site currently measures 0/100 — across 120 recorded answers, no engine cited us (we publish that). A print business we audited the same week measured 9/100: real presence, honestly marginal.

Flattering numbers are easy to sell once. Honest numbers are the only ones worth tracking over time — because when an honest number moves, something actually happened. That is the entire premise of how we measure.

Sources

methodology
ai-visibility
geo
measurement