MMatt Goren
← AI hub
GuideAI Search & AEOAEO & AI Search

How to Measure AEO: The Metrics That Actually Matter

Traffic is the wrong scoreboard for AI search. The number that matters is citation rate — the share of your real customer questions where an engine names you.

By Matt Goren · Updated July 29, 2026 · 6 min read

You cannot manage what you refuse to measure, and almost everyone measuring AI search is measuring the wrong thing. They open their analytics, look for a traffic spike, don't see one, and conclude AEO isn't working. Meanwhile an AI engine is naming their competitor to a buyer three times a day — a buyer who never clicks, never shows up in a report, and buys from someone else. The problem isn't that AEO can't be measured. It's that the scoreboard everyone is staring at was built for a different game.

The right metric is citation rate: the share of your real customer questions where an engine names you in its answer. Not traffic. Not rankings. Citation rate. I build and measure AEO for a living, and everything below is the method I actually use — how to build the test, run it across engines, and track the one number that tells you if you're winning. This is the measurement companion to the answer engine optimization playbook and tracking your AI visibility.

Start with why traffic lies to you

Classic analytics undercounts AI search by design, and you have to understand why before any of the numbers make sense. When someone asks ChatGPT "what's a good CRM for a two-person law firm" and it answers with three names including yours, that buyer got what they needed inside the answer. There was no click. No session in your analytics. No referrer. You were recommended to a qualified buyer and your reports show nothing.

This is the zero-click reality of AI search, and it means your traffic graph is blind to most of what AEO does. A flat referral line is not evidence AEO failed — it's evidence you're measuring the wrong surface. The answer is the product now, not the click, so the answer is what you have to measure directly.

Build a fixed prompt set

Measurement starts with a frozen list of the real questions your customers ask. This is the single most important setup step, and most people skip it.

Write down 30 to 50 buying questions — the specific, narrow ones a real customer asks before choosing, phrased the way they actually phrase them. Pull them from sales calls, support tickets, your on-site search bar, and the "how do I decide" moments your team hears every week. "Best project management tool" is too broad to be useful; "project management tool that works offline for a field crew" is the kind of question that actually gets asked and actually names a winner.

Then freeze the list. The power of a prompt set comes from running the exact same questions every time, so month-over-month results are comparable. Change the questions and you've thrown away your baseline. Write them once, lock them, and treat that list as your instrument.

Test across every engine, and record who got cited

Buyers don't all use one AI, so measure across all the ones that matter: ChatGPT, Claude, Perplexity, and Google's AI Overviews at a minimum. The same question can name completely different brands in different engines, and you need to see each surface honestly.

For every question in every engine, record one thing above all: were you named in the answer? Then capture the supporting detail — who else got cited, in what order, and why (a review, a comparison page, a spec the engine could quote). That "who else" column is where the real diagnosis lives, because it shows you exactly which competitor's content is beating yours and what the engine is pulling from them. My deeper breakdown of that decision is in how AI picks which brand to recommend.

One reading is a snapshot, not a measurement. Models are probabilistic — ask the same question twice and phrasing shifts — so run each prompt a couple of times and treat consistent citation, not a single lucky hit, as the real signal. And do it from a clean, logged-out session where you can, so your own history and location don't quietly personalize the answer and flatter you into thinking you're cited more than you are.

Track coverage over time

Turn all of that into one honest number: coverage rate — the percentage of your prompt set where you got cited. Fifty questions, cited in twelve, is 24% coverage. That's your baseline. It's blunt on purpose, because a blunt number you track every month beats a sophisticated one you check once.

Log it on a schedule. Monthly is right for most businesses — frequent enough to catch a drop, slow enough to avoid chasing noise. The value is entirely in the trend line: coverage climbing from 24% to 38% over a quarter is the clearest proof AEO is working that exists. Record each run in a simple sheet — date, engine, question, cited yes or no — so the history is there when you need to explain a jump or a dip. Add share of voice — your citations versus a named competitor's across the same set — and you can see whether you're gaining ground or just holding steady while someone else pulls ahead.

Watch leading indicators, not only lagging ones

Coverage rate is a lagging indicator — it tells you the outcome after the fact. To know where it's heading, watch the leading indicators, the inputs that cause citations:

  • Extractable-answer coverage — the share of your key pages that answer a real question in plain, quotable text up top.
  • Schema coverage — how much of your catalog carries clean, valid structured data, covered in schema and JSON-LD for AI search.
  • Crawler access — whether AI crawlers can actually reach and render your content, not just your human visitors.
  • Corroboration — fresh reviews, mentions, and comparisons across the web that back up your claims.

Leading indicators move first. When your extractable-answer and schema coverage climb, citation rate follows a few weeks later — which is exactly why you track both: inputs to know what to fix, outcomes to know if the fix worked. The full mechanics of building those inputs are in get cited by AI search.

What a good AEO dashboard tracks

Put it together and a real AEO dashboard has four things on it, in this order:

  1. Coverage rate over time — your headline number, one line, trending.
  2. Share of voice against your top competitors on the same prompt set.
  3. Per-engine breakdown — because winning in Perplexity and losing in Google AI is a fixable, specific problem.
  4. Leading indicators — extractable-answer coverage, schema coverage, crawler access — so you can see cause before effect.

Notice what's not the headline: raw traffic. Any referral clicks AI does send are a nice-to-have footnote, not the scoreboard. If you understand the difference between measuring AEO this way and the old SEO way, you understand the whole shift — I lay that out in AEO vs SEO.

Building that fixed prompt set, running it across four engines every month, and logging who got cited is real, repetitive work — and it only gets bigger as your question set and catalog grow. Doing the measurement and the fixing at scale is exactly the problem I built RunOctopus to solve: it runs the prompts, tracks your coverage across engines over time, and builds the extractable, cited-ready answer layer that moves the number — so measuring and improving AEO become one loop instead of two chores.

The businesses that win the next few years won't be the ones guessing whether AI mentions them. They'll be the ones who know their coverage rate to the point, watch it climb every month, and treat the answer — not the click — as the thing worth measuring.

#aeo#measurement#ai-search
Want to apply this right now?

Use the free, no-API prompt generators to put it into practice.

Open Prompt Studio →
Keep reading