How Arcalea measures AI visibility
The Arcalea AEO Index measures how visible a brand is inside AI answers. It scores two layers separately: the Memory layer, what AI models have internalized without web access, and the Retrieval layer, what they surface when grounded with live search. Each layer is scored on the same five signals, and the two are blended 0.32 Memory and 0.68 Retrieval into one index score. Separating the layers is what makes the score diagnostic rather than descriptive: a brand strong in Memory but weak in Retrieval is coasting on reputation, and a brand strong in Retrieval but weak in Memory is winning the live web without having entered the models' internal knowledge.
The architecture
What the AEO Index measures.
AI answer engines produce a recommendation in two steps. First they draw on what the model already knows, the knowledge compressed into its parameters during training. Then, when grounding is available, they search the live web and read what they find. Those two steps can disagree completely, and most measurement collapses them into a single number.
The Arcalea AEO Index does not collapse them. It runs the same prompt set twice, once with web access disabled and once with web access enabled, and scores each pass on its own.
Memory layer
What the model has internalized about the category, measured with web access disabled across 3 engines: ChatGPT, Claude, and Gemini. It reflects brand equity accumulated before the query, and it is slow to move.
Retrieval layer
What the model surfaces when grounded with live web search, measured across 5 engines: ChatGPT, Claude, Gemini, Perplexity, and Copilot. It reflects current citability, and it can move in weeks.
The scoring model
The five signals that determine AI visibility.
Both layers run the same five signals against two different sources of truth. Same arithmetic, different input.
Entity Mention Frequency (EMF)25%
Share of tested responses that mention the brand. Response count divided by total responses. The foundational frequency signal.
Topical Range Score20%
Fraction of distinct prompt categories in which the brand is mentioned. Independent of EMF: it measures breadth of presence, not volume of it.
Position Power Score20%
Average weighted position of the brand within an answer. Being named first carries more weight than being named last.
First-Position Rate15%
Fraction of mentions in which the brand is named first in a list. The advocacy signal, distinct from mere presence.
Cross-Platform Stability20%
One minus the coefficient of variation of per-platform mention counts. Evenness across engines, not just presence on them. A brand that dominates one engine and is absent from four scores poorly here, correctly.
Weights sum to 1.00. Each layer score is the weighted composite of its five signals.
Calculation
How the index score is calculated.
Each layer produces a composite from its five signals. The two layer composites are then blended.
Retrieval carries the larger weight for two reasons. Grounded search produces most of the answers that matter commercially. And retrieval responds to work in weeks, where memory responds over years.
The specific split is a standardization choice. Blend weights had been set per build, from 0.276 to 0.366, which left no two indexes on the same scale. We fixed the blend at the modal value across those builds so that they are. Anyone reproducing this should know the number was set for comparability across indexes, not optimized within one.
Worked example
Harvard Business School, from the August 2026 M7 Business Schools run.
| Signal | Weight | Memory | Retrieval |
|---|---|---|---|
| Entity Mention Frequency (EMF) | 0.25 | 0.673 | 0.555 |
| Topical Range Score | 0.20 | 1.000 | 1.000 |
| Position Power Score | 0.20 | 0.933 | 0.779 |
| First-Position Rate | 0.15 | 0.710 | 0.293 |
| Cross-Platform Stability | 0.20 | 0.897 | 0.838 |
| Layer composite | 0.841 | 0.706 |
Blended: (0.32 × 0.841) + (0.68 × 0.706) = 0.749
What the cells say that the headline cannot. Harvard's Memory score is the highest in the cohort. Its Retrieval score is fourth. The gap sits almost entirely in one cell: First-Position Rate falls from 0.710 in Memory to 0.293 in Retrieval. The models, drawing on what they know, name Harvard first in roughly seven of every ten answers that mention it. The models, reading the live web, in fewer than three. That is a specific, addressable content problem, and it is invisible in any single-number score.
Changelog
What we changed in v2, and why.
The index has run since early 2026. Version 2, published August 2026, makes two changes and retires one signal.
Change 1: the layers were separated and named. Earlier versions produced one score. v2 scores Memory and Retrieval independently on the same five signals and publishes both. The signal vocabulary did not change.
Change 2: the blend was standardized to 0.32 / 0.68 across every index. Blend weights were previously set per build, across a 0.276 to 0.366 range. Two indexes on different blends are not on the same scale, which made cross-index comparison invalid. 0.32 / 0.68 is now fixed in configuration for every vertical. Verified before the change rather than after: no headline cohort reordered in any index.
Retired: AI Share of Voice. In v2.5 the AI Share of Voice signal was removed from the composite because it correlated closely with Entity Mention Frequency. Both measure raw frequency, so carrying both gave frequency roughly 40% effective weight in a model that intended 25%. It was replaced by Topical Range Score, which measures breadth and is independent of EMF.
AI Share of Voice is still computed and still displayed. It remains a useful diagnostic. It is not part of the composite and does not appear in the weight table above. If you have seen an older Arcalea page describing AI Share of Voice as a 25% scoring weight, that page describes v1. This page is the methodology of record.
Scope and precision
What this index measures, and what it does not.
Four properties of how the index is built. Each one governs how a number should be read.
Cross-Platform Stability is defined within a layer. The Memory layer runs on 3 engines and the Retrieval layer on 5, because those are the engines that operate in each mode. The signal measures variance across per-platform counts, so it is scoped to its layer and always published with its platform count. Memory-Stability across 3 engines and Retrieval-Stability across 5 are each correct readings of their own layer.
Topical Range Score reports full coverage where full coverage exists. In the M7 Business Schools index it reads 1.000 for six of the seven schools. That is what a market leader looks like on this signal: present across every query category, full range, and the score says so. Where a category has spread, the signal spreads with it. Across the 50 entities in commercial debt collection its retrieval-layer values run 0.000 to 0.833.
Comparison is valid within an index and between its two layers. The blend is standardized across every index at 0.32 / 0.68. Prompt composition is not yet, so different categories ask structurally different questions and a 0.6 in one is not the same achievement as a 0.6 in another. Within a category, which is the unit that matters to a brand looking at its own market, comparison is exact. A common prompt set across categories is the next standard we are setting, and until it is in place we publish no cross-index rankings.
Entity resolution is maintained before anything is published. Organizations turn up in AI answers under acronyms, full names and bare domains, often all three in the same run. Those records are merged before publication, and citation domains tracked as sources are held separate from the brands under evaluation. Every leaderboard Arcalea publishes is deduplicated.
Collection
How the data is collected.
- Prompts are authored per category from real query patterns, covering recommendation, comparison, definitional, and specification intents.
- Each prompt runs against every engine in its layer, with repetitions, producing a response corpus scored per entity.
- The Memory pass runs with web access disabled. The Retrieval pass runs with search grounding enabled.
- Entities are tracked in tiers: primary, the operating companies in the category; intermediary, associations, trade press, and citation sources; extended, adjacent players. Tier is published alongside every entity, because mixing tiers in one ranking produces misleading comparisons.
- Scoring is deterministic. The same run scored twice produces identical results.
- Indexes are rerun quarterly and dated.
FAQ
Methodology FAQ.
How is the AEO Index different from a search ranking?
A search ranking measures position in a list of links. The AEO Index measures whether a brand is named inside a generated answer, how early it is named, how consistently it is named across engines, and across how many kinds of question. There is often no list of links to measure.
Why weight Retrieval higher than Memory?
Grounded search produces most of the answers that matter commercially, and retrieval responds to work in weeks where memory responds over years. Memory is not beyond your influence. It is just slow. You move it by building the kind of corroborated presence that future training runs absorb.
What does a high Memory score with a low Retrieval score mean?
The brand is coasting on reputation. AI models know it, but the live web is not confirming what they know. Since Retrieval carries 68% of the index score, that gap is expensive and it is fixable.
Can a brand score zero on the Memory layer?
Yes, and real companies do. A zero means the brand was not mentioned in any response, on any engine, with web access disabled. The data is there. The brand is not.
What is AI Share of Voice, and why is it not in the score?
It measures a brand's share of total category mentions. It was removed from the composite in v2.5 because it correlated with Entity Mention Frequency, and carrying both over-weighted raw frequency. It is still computed and displayed as a diagnostic.
How often is the index updated?
Quarterly, per category, with each edition named and dated.
See it applied
The indexes this methodology produces.
We publish the indexes rather than describing them. Each one carries the two-layer split, the per-signal detail, and the scoring model documented on this page.
Published benchmarks
- M7 Business Schools. The most competitive cohort we measure, and the source of the worked example above.
- Commercial Mechanical Contractors. A category where a standards body outranks most of the operating firms.
- Commercial Debt Collection. A category where the leader has no AI memory at all and wins on retrieval alone.
Also published
Higher Education. Surrogacy Agencies. Commercial Plumbing Distribution. Earlier builds in the research series, produced before the v2 two-layer model and scheduled to be rerun on the current standard.
Tracked but not published
We run this methodology across additional categories that we do not publish, for client-specific competitive analysis, and we build client-specific indexes on the same standard: the same two layers, the same five signals, the same blend. A client index is not a different product. It is this one, pointed at a category and an entity set that matter to one company.
If your category is not on the list above, that is what the index is for.
- Scoring model, signal definitions and blend weights are read directly from the Arcalea AEO Index scoring pipeline, which is the single source of truth for every figure on this page. The worked example uses the August 2026 M7 Business Schools run.
- Signal definitions are published as structured data using the schema.org DefinedTerm vocabulary, so they are machine-readable as well as human-readable. schema.org/DefinedTerm.
- Cross-Platform Stability is reported as one minus the coefficient of variation of per-platform mention counts. Coefficient of variation.
- Published sub-scores are computed values. No figure on this page is projected or illustrative.
Reviewed by Michael Stratta, Founder and CEO, Arcalea. Last updated August 2026.