Questions, not keywords

How do you build a buyer question set for AI visibility - and why do keyword lists fail?

A keyword list ranks a string. AI engines answer a question. If you're measuring the wrong instrument, every visibility number you calculate on top of it is noise wearing a decimal point.

28 Labs · September 2026

A buyer question set is a fixed list of full-sentence questions - the kind a real customer actually types into ChatGPT or asks Gemini - covering discovery, comparison, validation, and purchase stages, mostly unbranded, sized in the dozens, and frozen once you start measuring. It fails as a keyword list because AI engines don't rank a string; they read a full question, silently fan it into several searches of their own, and build the answer from what those searches return. Get the question set right and every number downstream means something. Get it lazy and the prettiest dashboard in the world is measuring nothing.

What makes a question "real"?

A real buyer question is a full sentence with context, not a fragment. "Best standing desk for a small home office under S$800" is a question. "Standing desk Singapore" is a keyword. Nobody talks to ChatGPT like a 2015 search bar - they type the way they'd ask a knowledgeable friend, with the budget, the location, the constraint, and the doubt built in.

That context isn't decoration. The words "under," "for," "near me," "worth it," and "vs" are what carry buyer intent. Strip them out to make a tidy keyword and you've thrown away the exact signal the AI engine uses to decide what the person actually needs. A question set has to preserve that context, question by question, or it isn't measuring what real buyers ask.

Cover the whole funnel, not just one stage

Visibility differs wildly by funnel stage, and most brands only build questions for one of them - usually comparison, because "X vs Y" is the easiest to write. A real set covers four:

A brand that's visible in comparison questions and invisible in discovery is losing every buyer before they've narrowed the field. A brand that only shows up at purchase never gets considered in the first place. You can't see either gap unless the set forces you to ask across all four stages.

The branded-question trapFill a question set with "tell me about BrandX" and you've built a flattery machine, not an instrument. AI engines default to describing a named brand favorably - that's not visibility, it's the engine being polite. Worse, no new customer has ever typed your brand name into ChatGPT before they've heard of you; branded questions test recall, not discovery. They belong in the set only as a small, deliberate slice - enough to check what the engine says when someone already knows your name - never as the bulk of it.

One market, sized right, frozen once set

Three more rules make the difference between an instrument and a vanity list.

1
One market, one language at a time
2
Draft mostly unbranded, full sentences
3
Balance across all four funnel stages
4
Freeze it, name the version, re-run identically
Change the questions and you've changed the ruler mid-measurement. Additions wait for the next named version.

Answers differ by market and language, so a question set is a per-market instrument, not a global one - the same brand can be visible in Dubai and invisible in London on the exact same category. Size matters too: enough questions for a rate to mean something statistically, few enough to hold steady wave over wave. We run 60 per market for that reason - fewer and one odd answer swings the whole number, more and you can't re-run the identical set reliably across four engines every wave.

Once a set is frozen, it stays frozen. Editing questions mid-measurement means you're no longer comparing wave to wave - you're comparing two different rulers and calling it a trend. New questions become a new version, dated and named, not a silent overwrite.

Why keyword lists fail

AI engines don't answer the question you asked - not directly. They read it, generate several reformulated searches behind the scenes, retrieve results for each, and merge everything into one answer. We cover the mechanics in what is query fan-out, but the short version: a single ChatGPT prompt can trigger anywhere from a handful to a dozen-plus internal searches, and Gemini tends to fan out even further. None of those internal searches are the keyword you'd have tracked in a rank tool.

Rank-tracking one string tells you where that string sits in a results list. It tells you nothing about who the engine actually names when it synthesizes an answer from a dozen sources - your page can rank for the fragment and still lose the sentence. That's the core mismatch: keyword tools measure search position, AI visibility depends on what gets picked and recommended once the fan-out has already happened. See how to navigate AEO for the broader shift this forces.

A keyword list looks like

  • "standing desk singapore"
  • "best standing desk"
  • "standing desk price"
  • One string, one intent guessed at

A question set looks like

  • "Best standing desk for a small HDB apartment under S$800?"
  • "Is a motorized desk worth it over a manual one?"
  • "Where can I buy one with fast delivery in Singapore?"
  • Full sentences, funnel stage, real constraint

A worked example: standing desks in Singapore

Illustrative only - not run data, just what a well-built slice looks like for a fictional category.

Eight questions, four stages, one market, one branded check at the end. Run this same eight against ChatGPT, Gemini, Claude, and Perplexity and you already know more about where a desk brand actually stands than a hundred-row keyword spreadsheet would tell you.

What we do about it

28 Labs builds a frozen, funnel-covering question set for each client's market before we run a single query. Every question is a full sentence a real buyer would ask, mapped to a funnel stage, weighted mostly unbranded, sized so the resulting rates hold statistical meaning, and versioned so nothing shifts mid-measurement. We ask that set across ChatGPT, Claude, Gemini, and Perplexity every wave, and every number we report traces back to a specific question and a specific answer. The set is the instrument. Build it honestly and the rest of the measurement takes care of itself.

28 Labs measures how AI engines answer real buyer questions - and what moves the counts. Every number we publish re-derives to specific questions and engines. try28labs.com