A four-layer model for AI discovery, and how to measure whether any of it worked.
Most AEO advice is an SEO checklist with the nouns swapped. The mechanism is genuinely different, and old habits break in four distinct places.
What follows isn't a list of tips. It's four layers. Eligibility - can engines reach and read you. Demand - which buyer decisions you're competing for. Evidence - what gives a buyer reason to choose you, on your pages and elsewhere. Measurement - whether they actually chose you more often.
| Layer | The question it answers |
|---|---|
| Eligibility | Can the engines reach and read you? |
| Demand | Which buyer decisions can you actually compete for? |
| Evidence | What gives an engine reason to recommend you? |
| Measurement | Did the answers actually change? |
They form a dependency stack rather than a workflow: weakness lower down constrains what the layers above can achieve. That isn't the same as a running order. You can research demand while a developer fixes access, and an unreadable site doesn't make third-party evidence worthless - engines routinely recommend a company entirely from sources it doesn't own. What a weak lower layer does is cap the ceiling of everything resting on it.
Answer-time fetchers read pages to build the answer a buyer is reading right now. Training crawlers collect content that may train future models. Blocking the second group is a legitimate business decision about your content, and it does not remove you from today's answers. As of August 2026 the answer-time agents are OAI-SearchBot and ChatGPT-User, Claude-SearchBot and Claude-User, PerplexityBot and Perplexity-User; the training crawlers include GPTBot and ClaudeBot. Google-Extended isn't a crawler at all - it has no user agent of its own and never fetches anything. These names change, so we keep a maintained version in our crawler guide rather than asking you to trust a blog post from last year. What doesn't change is the distinction, and getting it wrong in either direction is common and expensive.
The category has converged on deliverables that are cheap to produce and impossible to disprove. llms.txt is the clearest case: no AI engine has ever confirmed reading it, Google has said publicly that it doesn't support it and isn't planning to, and Ahrefs found 97% of these files across 137,000 domains received no requests at all. Watch for the laundering move when someone tells you otherwise - vendors cite AI companies "confirming support" when what those companies actually did was publish an llms.txt for their own documentation. Publishing isn't consuming. It costs almost nothing, so add it if you like; just never let it stand in for fetchable, readable, fast pages that answer something.
Keywords are fragments people type. The questions people ask an assistant are sentences, with context and constraints, and they don't come with a search volume figure you can look up. So the work starts by deciding which questions actually represent your buyers, in their vocabulary, for your market - and then holding that set still, so you can tell whether anything changed. No keyword tool contains "which mattress handles humidity best in Dubai".
Ask an engine "how do I choose a mattress" and you may get advice with no brand names in it at all. There's no position to win on that question this quarter - it's measuring the category, not you. Other questions route to a marketplace or a directory by default, where the reachable move is to be carried by that source rather than to outrank it. Three kinds of question, three different plans: the ones a brand can win, the ones that name nobody, and the ones held by an intermediary. Anyone promising movement before sorting your questions into those three piles is guessing.
Write pages that state facts a machine can lift: real prices, trial terms, addresses, timeframes, and who each option suits. An engine assembling an answer needs something it can quote and, ideally, corroborate elsewhere. A page that renders almost nothing without JavaScript offers neither, however good the copy is once it loads.
The logic on honesty is simpler than a study: "we're the best" gives an engine nothing to verify, while "they're stronger on price, we're stronger on materials and trial terms" does. We haven't isolated that effect in a controlled test and won't claim we have. Write like a source, not an ad.
This is the finding that should redirect most budgets. In Ahrefs' May 2026 study of 75,000 brands, YouTube mentions correlated with AI visibility at about 0.74 and branded web mentions at 0.66, while backlinks came in at 0.22.
Position seven on a search page still earns clicks, because the buyer sees ten options and picks one. An AI answer collapses the field into a shortlist, and can make an explicit recommendation. There is no position seven, and no reliable click curve running down the page - once a brand falls outside the shortlist the answer presents, its practical visibility collapses rather than tapering. That's why we count three separate things: whether you're named at all, whether you're inside the first three, and whether you're the single recommendation. The last is the closest answer-level analogue to winning the buyer's consideration - which is not the same as a sale, and we don't report it as one.
Engines change under you. A model update can stop an answer naming you with nothing having changed on your side, which means any before-and-after that doesn't record which model was running is telling you a story rather than a result. There's also no honest way to isolate one change when a client shipped six.
A frozen question set alone isn't enough either, because the same model can answer the same question differently twice. Freezing controls what you ask; repetition is what tells you whether an apparent movement survives repeated observation. The stronger design puts each question to each engine several times inside the same wave, so that variance can be observed rather than inferred afterwards. Where a measurement samples each question only once, that within-wave variance is simply invisible, and persistence across later waves is a weaker substitute - by then time and model state have moved as well. Worth being precise about that rather than implying it away. Either way, one answer flipping is not a finding, and any single observation is a draw from a distribution rather than a reading off a dial.
We run that test. What it shows, consistently, is that the more specific the position, the less stable it is: whether a brand is named at all holds up better across identical runs than whether it is in the top three, which holds up better than whether it is the single pick. The single recommendation is the noisiest thing on the board.
Which means a screenshot proving you appear for some question has a real chance of being an artifact of the moment it was asked, and anyone showing you a point-in-time win without a stability figure cannot tell you which it is. That includes us, unless we show you the spread.
One more thing about freezing: who freezes matters. If the party being measured also picks the questions, the panel drifts toward the questions that flatter. The set should be agreed with whoever the numbers will be reported to, before anything is asked, and then left alone.
And one limit worth stating plainly, because the category mostly doesn't: all of this measures retrieval - what engines fetch and say right now. What sits inside a model's training data is not directly observable, by us or by anyone. If a vendor offers to measure your training-data exposure, that isn't a capability that currently exists. We've written about where that line falls separately.
Start with eligibility, because poor first-party accessibility caps how much of the information environment you directly control. Then evidence on your own pages: make the buyer's question explicit, answer it immediately, and give specific facts a machine can extract without reconstructing your argument. Then corroboration off your site, which is where the strongest correlations live and which moves the slowest. Measurement runs throughout, not at the end.
Notice that the sequence and the strength ranking disagree, and that isn't a contradiction. Off-site presence shows the strongest correlation in the study above, and comes third. Highest leverage and first move are different questions.
What we would not do: publish at volume and hope, buy placements without weighing them as advertising, or promise a percentage lift from a mechanism nobody can verify. And we would not write the content and also operate the scoreboard that grades it. If the party producing your pages owns the number that says the pages worked, that number is marketing.
SEO gave marketers an observable chain: query, ranking, impression, click, session, conversion. AI answers break that chain. What you get instead is a question, a probabilistic answer, a recommendation, and then a gap where the attribution used to be.
The honest response is to treat this as an empirical discipline rather than a checklist. You change the information environment, put the questions to the engines, keep every answer they give, and check whether the answers moved. Freeze the buyer questions. Preserve the verbatim responses. Repeat enough to tell signal from variance. Narrow the attribution as far as the evidence honestly allows, and say so plainly when it won't narrow further.
The uncomfortable part of this category is that the honest version is slower and smaller than the pitch. It's also the only version you can audit. A score out of 100 you can't recompute isn't measurement.
More on how this works in practice, including the crawler detail and what we check when the numbers don't move: our guides.