Marketing Intelligence Insights and Research | Arcalea

How to Read an AEO Analysis | Arcalea

Written by Michael Stratta | Sep 8, 2026, 12:28:47 AM
Quick answer
Every AEO number is the last step in a chain of choices the analyst made and usually did not state. Two questions recover most of them: what was asked, and is there a margin of error? The answers tell you whether you are holding an anecdote, a snapshot, or a measurement. Each of those can support a different size of claim.

Two reports, one brand

Two reports about the same company land in the same month. One says the brand appears in 62% of ChatGPT answers. The other says 31%. The chief executive asks the two questions any executive would ask: which one is right, and can we move it?

Both reports are honest. Neither is wrong. They made different choices about how to ask, what to count, and how many times to look, and neither wrote those choices down. That is the normal condition of AEO reporting today, in software, consumer goods, financial services, healthcare, travel, and higher education alike. The example here is a mid-market software company. The same nine questions decode a report about a running shoe or a graduate program.

The two reports, decoded. Neither is wrong; the 31-point gap is the sum of choices neither report stated.

Three grades of evidence

An AEO analysis is one of three things, and all three are legitimate.

An anecdote is a screenshot, a single answer, or a handful of prompts with your name in them. It is good for spotting a problem and starting a conversation. It cannot tell you whether you moved, how you compare, or why anything happened.

A snapshot is dozens to a couple hundred prompts, on one AI product, on one day, often with competitors side by side. It is good for prioritizing: where to look first, a first ranking. It cannot tell you whether a change is real or noise, or whether anything you did caused it.

A measurement is many prompts, the same ones asked in every condition being compared, dated and repeated, with margins of error, your own earlier baseline, and claims limited to what was observed. It is good for tracking change and for testing whether an action moved the number. Even a measurement shows association rather than cause unless it was designed as an experiment, and its labels (what counted as a "mention") remain one team's judgment until someone else has checked a sample.

The one rule: match the claim to the grade. The mistake is never that someone took a screenshot. The mistake is asking a screenshot to carry a measurement's claim.

A pattern worth naming: the thinnest evidence tends to arrive with the biggest claims ("here is why you rank low, and here is the fix"), while the strongest evidence arrives with the most modest ones ("here is what changed, plus or minus this much"). When the explanation gets more confident and the evidence gets thinner, read more slowly.

Three grades, and the inversion: evidence strength rises left to right while the size of the "why" claim usually falls.

The two questions to ask first

What was asked?

Find the actual prompts. If the report says "prompt list available on request," request it. That footnote decides the grade. Prompts that contain your own name ("project tools like ours") pre-load the answer and flatter everyone. Generic buyer questions ("best project tool for a ten-person team") are what prospects actually type. A fixed, written-down set of generic questions that anyone could rerun is the mark of a measurement.

Two refinements. Removing your name is not enough on its own: a question that specifies a budget, a region, or a type of buyer can quietly favor one brand. And count questions rather than rewordings. A report that boasts "200 prompts" whose appendix shows 20 questions each phrased ten ways has asked 20 questions. Rewordings of one question form a question family, and a family counts once.

Is there a margin of error?

A margin of error is the range the true number plausibly sits in, given that only a sample of questions was asked. 6 mentions in 10 answers and 600 in 1,000 are both "60%," and they mean very different things; the first could easily be 40% or 80% next week. A report that shows only a point estimate ("visibility up 8 points since last quarter," with no range) leaves you unable to say whether 8 points is bigger than the noise. A range is the snapshot's mark. A range plus a statistical test (a check on whether a difference is larger than chance alone would produce) is the measurement's.

Both reports say "60%." With 10 answers the plausible range runs from 31% to 83%; with 1,000 it runs from 57% to 63%.

Apply both to the two reports. The 62% used branded prompts and shows no range. The 31% used generic buyer questions and shows a range of roughly 26% to 36%. Neither is wrong. They measured different things with different precision.

Seven more questions

Which AI, and in what setting?

"ChatGPT" is not one thing. Web search on or off, a personal account or a workplace account, memory on or off: each produces different answers. A report that names the product, the account type, the search mode, and the personalization settings is describing a repeatable condition. One that says "ChatGPT" is describing a brand of software. Watch especially for a single score labeled with one vendor's name that, on inspection, averages a consumer chat app and the same vendor's workplace product. One name, two products.

What was counted?

"Visibility" can mean four different things: your name appeared in a source link; your brand was mentioned; your brand was recommended as an option worth considering; your brand was recommended for the thing the question actually asked about. Those are separate counts. A brand recommended for a product it does not offer is a mention and a mismatch, and no report should score it as a win. A report that defined its metric before counting, and says what it excludes, has the measurement's habit.

How many, and what happened to the rest?

One answer is one draw. Dozens is a snapshot. Hundreds, the same hundreds in every condition compared so that each question is compared with itself, is a measurement. That pairing is what makes a comparison fair: only the questions whose answers flipped tell you anything. Then ask about the answers that did not count. A denominator is the number of completed answers a percentage is taken over; "31 of 48" is a denominator, and "31%" alone is not. A refusal ("I can't recommend specific products") is an answer and belongs in the count. A screenshot that cuts off before the answer ends is a missing answer, and a missing answer is never a zero.

Which raises the point most often missed: absence is not evidence of absence. Not being named in three answers is an anecdote about absence, with the same limits as an anecdote about presence. "We're invisible" from three screenshots is exactly as over-read as "we're winning" from three screenshots.

When?

AI answers change daily. An undated number is a photograph without a timestamp. A dated number from one day is a snapshot. A dated series is a measurement. If several AI products are compared, ask whether they were collected in the same window: five products collected weeks apart are also five dates.

Compared to what?

A bar chart of 20 competitors with percentages and no sample size or date is a snapshot until you know both. A table that places your number beside a competitor's number from a different report offers no comparison at all: different method, different day. The strongest comparison is your own baseline, your earlier number collected the same way, so that a change is a change and not a difference in method. External references (a published ranking, a market-share figure) are useful when labeled as what they are.

What is being claimed?

Read the headline against the grade. "Where you rank" is a snapshot's claim. "What changed, within this range" is a measurement's. "Why you rank where you do" is the hardest claim of all, and the one most often attached to the thinnest evidence. Association is not cause: two things moving together does not mean one produced the other.

Does it separate what was observed from what is recommended?

The most common sleight of hand in AEO reporting is a slide that moves from "you are under-mentioned" to a proposal in one step. A gap is observed, a cause is assumed, a solution is sold. Recommendations offered as hypotheses to test ("here is what we would try, and here is how you would know in ninety days whether it worked") mark a report you can act on. A recommendation is a hypothesis until a before-and-after measurement says otherwise.

If you get the analyst in the room

Five things you are unlikely to see on a slide but can ask about in the follow-up meeting.

  1. Who chose the questions, and were the sets kept apart? Buyer-style discovery questions, the questions your own team most wanted answered, and lookups of your brand by name are three legitimate sets that answer three different things. Ask which is which, and whether they were blended into one score.
  2. What happened to answers that did not count? Completed answers, refusals, and failed captures should each have a number.
  3. Were platforms compared at the same time? Differences between AI products collected on different dates are partly differences in dates.
  4. Did the results ever change the method? A product or setting should be swapped only for documented technical reasons, never because the first numbers disappointed.
  5. Who checked the coding, and could anyone re-check it? Deciding what counts as a mention is judgment. A team re-reading its own labels has not performed an independent check; the raw answers, with dates and settings, should exist so a second reader could look.

Decoding the two reports

The 62% report: branded prompts, no range, undated, no statement of settings. An anecdote, and a useful one, because it started the conversation and because its screenshots show what a prospect might actually see.

The 31% report: generic buyer questions, a stated range, one product with settings recorded, one collection date, competitors from the same run. A snapshot, and a good one. Neither report labeled who chose its questions: one asked what buyers in that market tend to ask, the other asked what the company most wanted to know. Both are fair questions. The 31-point gap between the reports is mostly the gap between those two question sets.

Neither can yet answer "can we move it?" That needs a measurement. The smallest upgrades that get there, in order of cost:

  1. Ask for the exact list of questions, with a note on who chose them. This one request settles branded versus generic, rewordings versus questions, and hidden favoritism in one stroke.
  2. Date it.
  3. Run the same questions again on a second date.
  4. Add a range and a denominator.
  5. Keep your own baseline, collected the same way, and compare to that.

The five-minute read

The decoder. Three forks place any AEO report in its grade; each grade names what it is good for and what it cannot carry.

Hold this next to the slide.

QuestionAnecdoteSnapshotMeasurement
Which AI, what setting?Not statedOne product, one settingEach product and setting labeled
What was counted?ImpressionsMentionsDefined before counting; mention, recommendation, and fit kept separate
What was asked?Branded promptsGeneric buyer questionsFixed, written-down set; who chose it labeled
How many, and the rest?OnceDozensHundreds, paired; refusals and missing counted
When?UndatedOne dayDated series
Margin of error?NoneA rangeA range and a test
Compared to what?NothingCompetitors, same runYour own baseline
What is claimed?Why you rankWhere you rankWhat changed, within limits
Observed vs. recommended?FusedFixes offeredHypotheses, with a test

Completed answers ___ · Refusals ___ · Missing ___

Tick your column. Then read the claims against it. A report whose headline can be restated as "we appeared in 31 of 48 answers to these 48 questions, on this product with these settings, on these dates" has told you its questions, its setting, its dates, and its denominator. If the headline cannot be restated that way from what is in the report, you have found what is missing.

Download the one-page checklist (PNG, prints in black and white).

Where measurement begins

The 62% slide is an anecdote, and anecdotes are where every good measurement program begins: someone noticed something and asked. The decoder's job is to keep that conversation moving toward measurement without dismissing where it started.

One thing to do tomorrow: ask for the question list.

Related reading