The build-or-buy of AI visibility

Should you buy an AI visibility tool or hire an agency?

Both sides of this market will tell you their option is the whole answer. It isn't. The question that actually decides it has nothing to do with features - it's who's going to read the numbers and do something about them.

28 Labs · September 2026

Buy a tool if you already have someone in-house who will open the dashboard every week, interpret what moved, and brief the fix work to your team. Hire an agency if you need someone else to design the questions, do that interpretation, and execute the fixes themselves. A dashboard nobody acts on is a subscription, not a strategy - and that's true whether the subscription is software or a services retainer.

What AI visibility tools do well

Self-serve and enterprise tracking tools are genuinely good at one thing: continuous, automated measurement at scale. They can run hundreds of prompts across multiple engines on a schedule, log every result, plot trend lines, and alert you when something moves. Per data point, they're cheap - a human running the same volume of queries by hand would cost far more than the software.

If your need is monitoring - "tell me when our visibility changes" - a tool answers that cleanly, and it does it without anyone on your side having to build the instrument first.

What tools don't do

A tool will track whatever questions you give it. It won't tell you if those are the right questions - garbage in still produces a clean-looking dashboard out. It won't tell you whether a drop is a model update, a competitor's move, or something on your own site, because that takes reading the actual answer text against what changed in the market that week. And it won't do the fix work: rewriting a page, restructuring an FAQ, closing a schema gap. It shows you the scoreboard. It doesn't coach the team.

That gap is exactly what what AI visibility tools can and can't measure goes through in more detail - the short version is that most tools measure presence, not why presence changed.

What agencies do well - and where to watch them

A good AEO agency earns its retainer on three things a tool can't do for itself: designing a question set that actually reflects how your buyers ask, interpreting what a movement means, and doing the execution - content, schema, structural fixes - that changes the answer. The agency also carries accountability. If visibility doesn't move, that's a conversation you can have with a person, not a support ticket.

Watch for two patterns that aren't worth the premium over software. The first is an agency that's really reselling a tool's dashboard with a markup and no added judgment - ask directly who designs your questions and whether you could re-derive every number yourself if you wanted to. The second is a content-volume shop rebadged as "AEO" - publishing pages at scale with no interpretation layer behind it. Neither is fraud, but neither is worth agency pricing either.

The decision factors

In-house capacity
Do you have someone who will own the tool's output every week - read it, interpret it, brief fixes? If the honest answer is no, a tool alone won't move anything.
Market and language complexity
One market, one language is easy to self-serve. Multiple markets, multiple languages, and the interpretation work multiplies faster than most in-house teams can absorb.
Monitoring vs. execution
If you just need to know where you stand, buy the instrument. If you need someone to change where you stand, you need execution capacity - either your own team or an agency's.
Budget shape
Software budget and services budget usually sit in different approval chains. Which one you actually have available can decide this before feature comparisons even start.
How contested your category is
A category where AI answers already name three competitors by default needs faster, more deliberate intervention than one where nobody's shown up yet.

The comparison, plainly

A tool is built for

  • Continuous tracking across many prompts and engines
  • Dashboards, trend lines, alerting
  • Scale at low cost per data point
  • Teams that already have an owner for the output

An agency is built for

  • Designing the question set that actually matters
  • Interpreting what moved and why
  • Doing the fix work itself
  • Being accountable for the outcome, not just the reading
Self-serve tracking tools run from the tens to a few hundred US dollars a month. Enterprise platforms with more engines, markets, and seats generally start around $500 a month and climb past $3,000. AEO agency retainers in the UAE typically run roughly AED 4,000 to AED 20,000 a month, depending on market count and how much of the retainer is measurement versus execution.

The hybrid most mature teams end up running

In practice, the false choice resolves itself once a program matures. A tool handles continuous monitoring across the long tail of prompts. A specialist handles the waves that matter - the quarterly deep read, the interpretation of what actually shifted, and the execution once you know what to fix. Neither replaces the other; they cover different parts of the same job.

The two failure modes we see most often sit on opposite ends of this. The first is buying the tool, getting the dashboard live, and discovering three months later that nobody on the team ever opens it - that's the most common way this money gets wasted. The second is paying an agency for "measurement" that turns out to be the same tool's numbers with a summary email attached. Ask who designs the questions. Ask if you could re-derive the numbers yourself. If pricing feels opaque relative to what's actually delivered, AEO retainer pricing in the UAE breaks down what a realistic range buys at each tier.

The honest readThis isn't tools versus agencies - both are right for different buyers, and plenty of good programs run both at once. The decision that actually predicts whether you'll see results is who's going to act on the numbers once you have them. Buy or hire around that answer, not around feature lists.

What we do about it

We don't sell a dashboard on its own. We design the question set from real buyer language, run it across ChatGPT, Claude, Gemini, and Perplexity on a wave schedule, and every count we publish re-derives to a specific question and a specific answer - so you can check our work, not just trust it. Where a client already runs a tracking tool, we treat its output as one more input to the interpretation, not a replacement for it. The instrument matters less than whether someone reads what it says and acts.

28 Labs measures how AI engines answer real buyer questions - and what moves the counts. Every number we publish re-derives to specific questions and engines. try28labs.com