A keyword list ranks a string. AI engines answer a question. If you're measuring the wrong instrument, every visibility number you calculate on top of it is noise wearing a decimal point.
A real buyer question is a full sentence with context, not a fragment. "Best standing desk for a small home office under S$800" is a question. "Standing desk Singapore" is a keyword. Nobody talks to ChatGPT like a 2015 search bar - they type the way they'd ask a knowledgeable friend, with the budget, the location, the constraint, and the doubt built in.
That context isn't decoration. The words "under," "for," "near me," "worth it," and "vs" are what carry buyer intent. Strip them out to make a tidy keyword and you've thrown away the exact signal the AI engine uses to decide what the person actually needs. A question set has to preserve that context, question by question, or it isn't measuring what real buyers ask.
Visibility differs wildly by funnel stage, and most brands only build questions for one of them - usually comparison, because "X vs Y" is the easiest to write. A real set covers four:
A brand that's visible in comparison questions and invisible in discovery is losing every buyer before they've narrowed the field. A brand that only shows up at purchase never gets considered in the first place. You can't see either gap unless the set forces you to ask across all four stages.
Three more rules make the difference between an instrument and a vanity list.
Answers differ by market and language, so a question set is a per-market instrument, not a global one - the same brand can be visible in Dubai and invisible in London on the exact same category. Size matters too: enough questions for a rate to mean something statistically, few enough to hold steady wave over wave. We run 60 per market for that reason - fewer and one odd answer swings the whole number, more and you can't re-run the identical set reliably across four engines every wave.
Once a set is frozen, it stays frozen. Editing questions mid-measurement means you're no longer comparing wave to wave - you're comparing two different rulers and calling it a trend. New questions become a new version, dated and named, not a silent overwrite.
AI engines don't answer the question you asked - not directly. They read it, generate several reformulated searches behind the scenes, retrieve results for each, and merge everything into one answer. We cover the mechanics in what is query fan-out, but the short version: a single ChatGPT prompt can trigger anywhere from a handful to a dozen-plus internal searches, and Gemini tends to fan out even further. None of those internal searches are the keyword you'd have tracked in a rank tool.
Rank-tracking one string tells you where that string sits in a results list. It tells you nothing about who the engine actually names when it synthesizes an answer from a dozen sources - your page can rank for the fragment and still lose the sentence. That's the core mismatch: keyword tools measure search position, AI visibility depends on what gets picked and recommended once the fan-out has already happened. See how to navigate AEO for the broader shift this forces.
Illustrative only - not run data, just what a well-built slice looks like for a fictional category.
Eight questions, four stages, one market, one branded check at the end. Run this same eight against ChatGPT, Gemini, Claude, and Perplexity and you already know more about where a desk brand actually stands than a hundred-row keyword spreadsheet would tell you.
28 Labs builds a frozen, funnel-covering question set for each client's market before we run a single query. Every question is a full sentence a real buyer would ask, mapped to a funnel stage, weighted mostly unbranded, sized so the resulting rates hold statistical meaning, and versioned so nothing shifts mid-measurement. We ask that set across ChatGPT, Claude, Gemini, and Perplexity every wave, and every number we report traces back to a specific question and a specific answer. The set is the instrument. Build it honestly and the rest of the measurement takes care of itself.