The short version: tools can measure what AI engines answer today. They can't see inside the model. Anyone who says otherwise is selling you something.
Step 2 is where measurement lives. When an engine searches the web to build an answer, the pages it reads and the brands it names are observable facts - ask the same questions on a schedule and you get a real instrument: who's named, who's recommended, which pages the answers came from.
Step 3 is where the overclaiming lives. Models also carry what they absorbed in training - and there's no API for that. You can see its fingerprints when an engine names a brand without reading any page. But no tool can audit it, and nobody can promise to change it on a deadline.
Retrieval is fixable on a timeline you control: publish the page that answers the question, get into the sources engines read, and the counts can move inside weeks. Training memory moves slowly and indirectly - it follows what the public web says about you, over months. So an honest engagement measures retrieval, fixes retrieval, and treats the slow layer as a tailwind you earn - not a dial anyone can turn.
We measure the retrieval layer against a frozen set of real buyer questions - the same questions every wave, so movement means something. We name this limit in every report we ship, because clients who understand the instrument make better decisions with it. That's the whole trick: measure what's measurable, be honest about the rest, and let the counts speak.
Related: Microsoft now publishes first-party citation data to site owners - see what Bing's AI Performance report can and can't tell you.