Two reports about the same company land in the same month. One says the brand appears in 62% of ChatGPT answers. The other says 31%. The chief executive asks the two questions any executive would ask: which one is right, and can we move it?
Both reports are honest. Neither is wrong. They made different choices about how to ask, what to count, and how many times to look, and neither wrote those choices down. That is the normal condition of AEO reporting today, in software, consumer goods, financial services, healthcare, travel, and higher education alike. The example here is a mid-market software company. The same nine questions decode a report about a running shoe or a graduate program.
An AEO analysis is one of three things, and all three are legitimate.
An anecdote is a screenshot, a single answer, or a handful of prompts with your name in them. It is good for spotting a problem and starting a conversation. It cannot tell you whether you moved, how you compare, or why anything happened.
A snapshot is dozens to a couple hundred prompts, on one AI product, on one day, often with competitors side by side. It is good for prioritizing: where to look first, a first ranking. It cannot tell you whether a change is real or noise, or whether anything you did caused it.
A measurement is many prompts, the same ones asked in every condition being compared, dated and repeated, with margins of error, your own earlier baseline, and claims limited to what was observed. It is good for tracking change and for testing whether an action moved the number. Even a measurement shows association rather than cause unless it was designed as an experiment, and its labels (what counted as a "mention") remain one team's judgment until someone else has checked a sample.
The one rule: match the claim to the grade. The mistake is never that someone took a screenshot. The mistake is asking a screenshot to carry a measurement's claim.
A pattern worth naming: the thinnest evidence tends to arrive with the biggest claims ("here is why you rank low, and here is the fix"), while the strongest evidence arrives with the most modest ones ("here is what changed, plus or minus this much"). When the explanation gets more confident and the evidence gets thinner, read more slowly.
Find the actual prompts. If the report says "prompt list available on request," request it. That footnote decides the grade. Prompts that contain your own name ("project tools like ours") pre-load the answer and flatter everyone. Generic buyer questions ("best project tool for a ten-person team") are what prospects actually type. A fixed, written-down set of generic questions that anyone could rerun is the mark of a measurement.
Two refinements. Removing your name is not enough on its own: a question that specifies a budget, a region, or a type of buyer can quietly favor one brand. And count questions rather than rewordings. A report that boasts "200 prompts" whose appendix shows 20 questions each phrased ten ways has asked 20 questions. Rewordings of one question form a question family, and a family counts once.
A margin of error is the range the true number plausibly sits in, given that only a sample of questions was asked. 6 mentions in 10 answers and 600 in 1,000 are both "60%," and they mean very different things; the first could easily be 40% or 80% next week. A report that shows only a point estimate ("visibility up 8 points since last quarter," with no range) leaves you unable to say whether 8 points is bigger than the noise. A range is the snapshot's mark. A range plus a statistical test (a check on whether a difference is larger than chance alone would produce) is the measurement's.
Apply both to the two reports. The 62% used branded prompts and shows no range. The 31% used generic buyer questions and shows a range of roughly 26% to 36%. Neither is wrong. They measured different things with different precision.
"ChatGPT" is not one thing. Web search on or off, a personal account or a workplace account, memory on or off: each produces different answers. A report that names the product, the account type, the search mode, and the personalization settings is describing a repeatable condition. One that says "ChatGPT" is describing a brand of software. Watch especially for a single score labeled with one vendor's name that, on inspection, averages a consumer chat app and the same vendor's workplace product. One name, two products.
"Visibility" can mean four different things: your name appeared in a source link; your brand was mentioned; your brand was recommended as an option worth considering; your brand was recommended for the thing the question actually asked about. Those are separate counts. A brand recommended for a product it does not offer is a mention and a mismatch, and no report should score it as a win. A report that defined its metric before counting, and says what it excludes, has the measurement's habit.
One answer is one draw. Dozens is a snapshot. Hundreds, the same hundreds in every condition compared so that each question is compared with itself, is a measurement. That pairing is what makes a comparison fair: only the questions whose answers flipped tell you anything. Then ask about the answers that did not count. A denominator is the number of completed answers a percentage is taken over; "31 of 48" is a denominator, and "31%" alone is not. A refusal ("I can't recommend specific products") is an answer and belongs in the count. A screenshot that cuts off before the answer ends is a missing answer, and a missing answer is never a zero.
Which raises the point most often missed: absence is not evidence of absence. Not being named in three answers is an anecdote about absence, with the same limits as an anecdote about presence. "We're invisible" from three screenshots is exactly as over-read as "we're winning" from three screenshots.
AI answers change daily. An undated number is a photograph without a timestamp. A dated number from one day is a snapshot. A dated series is a measurement. If several AI products are compared, ask whether they were collected in the same window: five products collected weeks apart are also five dates.
A bar chart of 20 competitors with percentages and no sample size or date is a snapshot until you know both. A table that places your number beside a competitor's number from a different report offers no comparison at all: different method, different day. The strongest comparison is your own baseline, your earlier number collected the same way, so that a change is a change and not a difference in method. External references (a published ranking, a market-share figure) are useful when labeled as what they are.
Read the headline against the grade. "Where you rank" is a snapshot's claim. "What changed, within this range" is a measurement's. "Why you rank where you do" is the hardest claim of all, and the one most often attached to the thinnest evidence. Association is not cause: two things moving together does not mean one produced the other.
The most common sleight of hand in AEO reporting is a slide that moves from "you are under-mentioned" to a proposal in one step. A gap is observed, a cause is assumed, a solution is sold. Recommendations offered as hypotheses to test ("here is what we would try, and here is how you would know in ninety days whether it worked") mark a report you can act on. A recommendation is a hypothesis until a before-and-after measurement says otherwise.
Five things you are unlikely to see on a slide but can ask about in the follow-up meeting.
The 62% report: branded prompts, no range, undated, no statement of settings. An anecdote, and a useful one, because it started the conversation and because its screenshots show what a prospect might actually see.
The 31% report: generic buyer questions, a stated range, one product with settings recorded, one collection date, competitors from the same run. A snapshot, and a good one. Neither report labeled who chose its questions: one asked what buyers in that market tend to ask, the other asked what the company most wanted to know. Both are fair questions. The 31-point gap between the reports is mostly the gap between those two question sets.
Neither can yet answer "can we move it?" That needs a measurement. The smallest upgrades that get there, in order of cost:
Hold this next to the slide.
| Question | Anecdote | Snapshot | Measurement |
|---|---|---|---|
| Which AI, what setting? | Not stated | One product, one setting | Each product and setting labeled |
| What was counted? | Impressions | Mentions | Defined before counting; mention, recommendation, and fit kept separate |
| What was asked? | Branded prompts | Generic buyer questions | Fixed, written-down set; who chose it labeled |
| How many, and the rest? | Once | Dozens | Hundreds, paired; refusals and missing counted |
| When? | Undated | One day | Dated series |
| Margin of error? | None | A range | A range and a test |
| Compared to what? | Nothing | Competitors, same run | Your own baseline |
| What is claimed? | Why you rank | Where you rank | What changed, within limits |
| Observed vs. recommended? | Fused | Fixes offered | Hypotheses, with a test |
Completed answers ___ · Refusals ___ · Missing ___
Tick your column. Then read the claims against it. A report whose headline can be restated as "we appeared in 31 of 48 answers to these 48 questions, on this product with these settings, on these dates" has told you its questions, its setting, its dates, and its denominator. If the headline cannot be restated that way from what is in the report, you have found what is missing.
Download the one-page checklist (PNG, prints in black and white).
The 62% slide is an anecdote, and anecdotes are where every good measurement program begins: someone noticed something and asked. The decoder's job is to keep that conversation moving toward measurement without dismissing where it started.
One thing to do tomorrow: ask for the question list.