Methodology

How Arcalea measures AI visibility

Two layers, five signals, published weights, a worked example, and the exact scope of what the numbers support. This is the methodology of record for every Arcalea AEO Index.
MS
Michael Stratta
Founder & CEO, Arcalea
Version 2 · Updated August 2026
Quick answer

The Arcalea AEO Index measures how visible a brand is inside AI answers. It scores two layers separately: the Memory layer, what AI models have internalized without web access, and the Retrieval layer, what they surface when grounded with live search. Each layer is scored on the same five signals, and the two are blended 0.32 Memory and 0.68 Retrieval into one index score. Separating the layers is what makes the score diagnostic rather than descriptive: a brand strong in Memory but weak in Retrieval is coasting on reputation, and a brand strong in Retrieval but weak in Memory is winning the live web without having entered the models' internal knowledge.

The architecture

What the AEO Index measures.

AI answer engines produce a recommendation in two steps. First they draw on what the model already knows, the knowledge compressed into its parameters during training. Then, when grounding is available, they search the live web and read what they find. Those two steps can disagree completely, and most measurement collapses them into a single number.

The Arcalea AEO Index does not collapse them. It runs the same prompt set twice, once with web access disabled and once with web access enabled, and scores each pass on its own.

Memory layer

What the model has internalized about the category, measured with web access disabled across 3 engines: ChatGPT, Claude, and Gemini. It reflects brand equity accumulated before the query, and it is slow to move.

Retrieval layer

What the model surfaces when grounded with live web search, measured across 5 engines: ChatGPT, Claude, Gemini, Perplexity, and Copilot. It reflects current citability, and it can move in weeks.

The scoring model

The five signals that determine AI visibility.

Both layers run the same five signals against two different sources of truth. Same arithmetic, different input.

Entity Mention Frequency (EMF)25%

Share of tested responses that mention the brand. Response count divided by total responses. The foundational frequency signal.

Topical Range Score20%

Fraction of distinct prompt categories in which the brand is mentioned. Independent of EMF: it measures breadth of presence, not volume of it.

Position Power Score20%

Average weighted position of the brand within an answer. Being named first carries more weight than being named last.

First-Position Rate15%

Fraction of mentions in which the brand is named first in a list. The advocacy signal, distinct from mere presence.

Cross-Platform Stability20%

One minus the coefficient of variation of per-platform mention counts. Evenness across engines, not just presence on them. A brand that dominates one engine and is absent from four scores poorly here, correctly.

Weights sum to 1.00. Each layer score is the weighted composite of its five signals.

Calculation

How the index score is calculated.

Each layer produces a composite from its five signals. The two layer composites are then blended.

Index score = (0.32 × Memory) + (0.68 × Retrieval)

Retrieval carries the larger weight for two reasons. Grounded search produces most of the answers that matter commercially. And retrieval responds to work in weeks, where memory responds over years.

The specific split is a standardization choice. Blend weights had been set per build, from 0.276 to 0.366, which left no two indexes on the same scale. We fixed the blend at the modal value across those builds so that they are. Anyone reproducing this should know the number was set for comparability across indexes, not optimized within one.

Worked example

Harvard Business School, from the August 2026 M7 Business Schools run.

The 2x5 model, worked on one brand Five weighted signals scored twice: once in the Memory layer across three engines with web access disabled, once in the Retrieval layer across five engines with live search. Harvard Business School scores 0.841 on Memory and 0.706 on Retrieval, blending at 0.32 and 0.68 to an index score of 0.749. The gap sits almost entirely in First-Position Rate, which falls from 0.710 to 0.293. The same five signals, measured twice WORKED ON HARVARD BUSINESS SCHOOL, M7 INDEX, AUGUST 2026 Memory layer 3 engines, no web access Retrieval layer 5 engines, live search THE FIVE SIGNALS Entity Mention Frequency Share of responses that name you 25% 0.673 0.555 Topical Range Score How many question types you cover 20% 1.000 1.000 Position Power Score How early you appear in the answer 20% 0.933 0.779 First-Position Rate How often you are named first 15% 0.710 0.293 Cross-Platform Stability How evenly you show up across engines 20% 0.897 0.838 Each layer is the weighted composite of its five signals. 0.841 0.706 (0.32 x 0.841) + (0.68 x 0.706) = 0.749 The gap is one cell. Known first, found first far less often.
Five signals, scored twice, blended once. The gap between what models know and what they find sits in a single cell.
SignalWeightMemoryRetrieval
Entity Mention Frequency (EMF)0.250.6730.555
Topical Range Score0.201.0001.000
Position Power Score0.200.9330.779
First-Position Rate0.150.7100.293
Cross-Platform Stability0.200.8970.838
Layer composite0.8410.706

Blended: (0.32 × 0.841) + (0.68 × 0.706) = 0.749

What the cells say that the headline cannot. Harvard's Memory score is the highest in the cohort. Its Retrieval score is fourth. The gap sits almost entirely in one cell: First-Position Rate falls from 0.710 in Memory to 0.293 in Retrieval. The models, drawing on what they know, name Harvard first in roughly seven of every ten answers that mention it. The models, reading the live web, in fewer than three. That is a specific, addressable content problem, and it is invisible in any single-number score.

Changelog

What we changed in v2, and why.

The index has run since early 2026. Version 2, published August 2026, makes two changes and retires one signal.

Change 1: the layers were separated and named. Earlier versions produced one score. v2 scores Memory and Retrieval independently on the same five signals and publishes both. The signal vocabulary did not change.

Change 2: the blend was standardized to 0.32 / 0.68 across every index. Blend weights were previously set per build, across a 0.276 to 0.366 range. Two indexes on different blends are not on the same scale, which made cross-index comparison invalid. 0.32 / 0.68 is now fixed in configuration for every vertical. Verified before the change rather than after: no headline cohort reordered in any index.

Retired: AI Share of Voice. In v2.5 the AI Share of Voice signal was removed from the composite because it correlated closely with Entity Mention Frequency. Both measure raw frequency, so carrying both gave frequency roughly 40% effective weight in a model that intended 25%. It was replaced by Topical Range Score, which measures breadth and is independent of EMF.

AI Share of Voice is still computed and still displayed. It remains a useful diagnostic. It is not part of the composite and does not appear in the weight table above. If you have seen an older Arcalea page describing AI Share of Voice as a 25% scoring weight, that page describes v1. This page is the methodology of record.

Scope and precision

What this index measures, and what it does not.

Four properties of how the index is built. Each one governs how a number should be read.

Cross-Platform Stability is defined within a layer. The Memory layer runs on 3 engines and the Retrieval layer on 5, because those are the engines that operate in each mode. The signal measures variance across per-platform counts, so it is scoped to its layer and always published with its platform count. Memory-Stability across 3 engines and Retrieval-Stability across 5 are each correct readings of their own layer.

Topical Range Score reports full coverage where full coverage exists. In the M7 Business Schools index it reads 1.000 for six of the seven schools. That is what a market leader looks like on this signal: present across every query category, full range, and the score says so. Where a category has spread, the signal spreads with it. Across the 50 entities in commercial debt collection its retrieval-layer values run 0.000 to 0.833.

Comparison is valid within an index and between its two layers. The blend is standardized across every index at 0.32 / 0.68. Prompt composition is not yet, so different categories ask structurally different questions and a 0.6 in one is not the same achievement as a 0.6 in another. Within a category, which is the unit that matters to a brand looking at its own market, comparison is exact. A common prompt set across categories is the next standard we are setting, and until it is in place we publish no cross-index rankings.

Entity resolution is maintained before anything is published. Organizations turn up in AI answers under acronyms, full names and bare domains, often all three in the same run. Those records are merged before publication, and citation domains tracked as sources are held separate from the brands under evaluation. Every leaderboard Arcalea publishes is deduplicated.

Collection

How the data is collected.

  • Prompts are authored per category from real query patterns, covering recommendation, comparison, definitional, and specification intents.
  • Each prompt runs against every engine in its layer, with repetitions, producing a response corpus scored per entity.
  • The Memory pass runs with web access disabled. The Retrieval pass runs with search grounding enabled.
  • Entities are tracked in tiers: primary, the operating companies in the category; intermediary, associations, trade press, and citation sources; extended, adjacent players. Tier is published alongside every entity, because mixing tiers in one ranking produces misleading comparisons.
  • Scoring is deterministic. The same run scored twice produces identical results.
  • Indexes are rerun quarterly and dated.

FAQ

Methodology FAQ.

How is the AEO Index different from a search ranking?

A search ranking measures position in a list of links. The AEO Index measures whether a brand is named inside a generated answer, how early it is named, how consistently it is named across engines, and across how many kinds of question. There is often no list of links to measure.

Why weight Retrieval higher than Memory?

Grounded search produces most of the answers that matter commercially, and retrieval responds to work in weeks where memory responds over years. Memory is not beyond your influence. It is just slow. You move it by building the kind of corroborated presence that future training runs absorb.

What does a high Memory score with a low Retrieval score mean?

The brand is coasting on reputation. AI models know it, but the live web is not confirming what they know. Since Retrieval carries 68% of the index score, that gap is expensive and it is fixable.

Can a brand score zero on the Memory layer?

Yes, and real companies do. A zero means the brand was not mentioned in any response, on any engine, with web access disabled. The data is there. The brand is not.

What is AI Share of Voice, and why is it not in the score?

It measures a brand's share of total category mentions. It was removed from the composite in v2.5 because it correlated with Entity Mention Frequency, and carrying both over-weighted raw frequency. It is still computed and displayed as a diagnostic.

How often is the index updated?

Quarterly, per category, with each edition named and dated.