Ask ChatGPT the same question twice and you can get two different lists of brands. Ask it again next month and the model itself might not be the one you asked the first time. None of that is a glitch.
Large language models don't look answers up. They generate them, one token at a time, by sampling from a distribution of likely next words. That's nondeterminism, and it's not a bug someone forgot to fix - it's how the technology produces language at all.
Send the identical prompt to the identical engine twice in a row and the brand list, the order, even the wording can shift. Nobody touched anything. If your only evidence is one chat window, you've sampled the distribution once and mistaken it for the truth.
Most engines with browsing pull live pages to ground their answers. That index is not static. A competitor publishes a new comparison page, a listicle gets updated, your own page changes or drops out of the crawl - and the set of sources the engine reads shifts under it. The model didn't change. The web it's reading did.
This is retrieval freshness, and it means your visibility can move for reasons that have nothing to do with your product, your content, or your competitors doing anything unusual - just the ordinary churn of the web getting re-crawled and re-ranked.
Underneath the chat window, the model is not fixed. OpenAI shipped GPT-5.4 on March 5, 2026, then GPT-5.5 on April 23, 2026 - what it called the biggest jump since GPT-5 launched in August 2025. GPT-5.1 was pulled from ChatGPT on March 11, 2026; GPT-5.2 followed on June 12, 2026, with existing conversations moved onto GPT-5.5 automatically. That's roughly a model swap every six to eight weeks, each one capable of reshuffling how an entire category gets answered overnight, with no announcement aimed at any single brand.
Google, Anthropic, and Perplexity run their own release cadences behind the same kind of chat window. You will rarely be told when an update lands. You'll just see the answers move.
The fourth source of change isn't the model or the sources - it's the format. OpenAI began testing ads inside free ChatGPT conversations in 2026. Google's AI Overviews add and drop shopping carousels, maps, and follow-up prompts depending on the query. When the surface changes, the thing you're trying to measure - "did we get mentioned" - can change shape entirely, independent of anything the model decided about your brand.
Picture the failure mode: the founder opens ChatGPT, types the question once, and screenshots what comes back. If the brand shows up, the screenshot goes in the deck as proof the AEO work is paying off. If it doesn't, it goes in a panicked Slack message asking why the agency isn't doing anything.
Both reactions are wrong for the same reason. One check can't distinguish any of the four causes above from each other, and it ages instantly in both directions - the good screenshot is false comfort by the time anyone reads the deck, and the bad one is often just noise from an engine that will answer differently five minutes later. Neither tells you anything you can act on.
The boardroom version of this is worse. An exec sees one bad answer in a live demo, someone promises the next one will be clean, and nobody in the room ever defined what "clean" means as a repeatable count. A promise with no denominator isn't a target. It's a hope.
You can't eliminate the four sources of change above. Nobody can - not us, not the engines themselves. What you can do is stop asking a single question once and start asking a frozen set of real buyer questions, repeated on a fixed schedule, across multiple engines, so movement can be separated from noise.
That's the discipline behind what we call a wave: the same question set, asked the same way, on a schedule, with every count re-derivable back to a specific question and a specific answer. When a number moves between waves, you can check which of the four causes is the likely source - and sometimes the honest answer is that the swing is the engine's doing, not yours or your agency's. We cover what that separation looks like in practice in when AI visibility numbers don't move, and where the limits of any measurement approach sit in what AI visibility tools can and can't measure.
We run a frozen question set across ChatGPT, Claude, Gemini, and Perplexity on a repeating wave schedule, not a one-off check. Every count we report re-derives to specific questions and specific answers, so when something moves, we can tell you whether it's a real shift in how you're being recommended or one of the four sources of noise above. Volatility is exactly why this measurement exists - in a system this noisy, a repeated count against a frozen denominator is the only thing that means anything.