MBA-AEPI · Wave 1 · September 2026 · Public Release

The complete AI-visibility index for MBA programs.

Two modes of AI visibility: what ChatGPT already knows from training, and what it says once it can search live. Measured separately, scored honestly, and compared in one view. Wave 1: ChatGPT GPT-4o.

25 MBA programs 600 dual-mode prompts Memory + live retrieval Two scores, honestly separated
New instrument This is Wave 1 of MBA-AEPI, a companion measurement to the AEO Index: M7 Business Schools. It uses a different methodology, a single-platform (ChatGPT) baseline, and a broader 25-program cohort. It does not replace the flagship M7 Index; see that page for the full 5-platform, memory-vs-retrieval measurement.
25
Programs tracked
Stanford 0.627
LSS leader
0.273
Median M₁ᴬᴳˢᴼ
0%
Negative valence
25 Stable
ΔM₁ classification
1,200
Total captures
100%
Completion rate
Wave 1 What AI already says about MBA programs
The short version

Executive Summary

Everything this index measures, in ninety seconds. Every figure is drawn from the panels below, so the summary and the evidence cannot disagree.

0.627
Stanford GSB: #1 M₁ᴬᴳˢᴼ
Appears in 62.7% of all live-search responses across 600 prompts
25
Programs, all Stable
Zero Search-Amplified or Search-Suppressed at Wave 1 n=600
0%
Negative valence
ChatGPT acts as a recommender, with no critical framing detected
0.63
MIT Sloan: top ECM
Most consistent mention rate across all four prompt strata
Finding Stanford GSB, Wharton, and Harvard Business School lead the AI-surface contest and are statistically tied: their Wilson 95% CIs overlap. Below Booth (rank 8, 0.442), visibility drops sharply. 17 of 25 programs score below 0.10, meaning ChatGPT mentions them in fewer than 1 in 10 responses. For those programs, AI visibility is an unaddressed competitive risk.
What this means M₁ᴬᴳˢᴼ is not a prestige index. It measures AI-surface salience, which diverges from US News rankings in measurable ways. Columbia Business School outscores Kellogg in LSS mode despite similar traditional rankings. MIT Sloan's ECM of 0.63 reflects cross-stratum robustness unavailable to schools with equal headline scores but narrow visibility profiles.
Act 1 Where programs stand
Primary metric

The M₁ᴬᴳˢᴼ headline index

M₁ᴬᴳˢᴼ = proportion of ChatGPT Live Search Surface responses in which each program is mentioned. n=600 prompts. Wilson 95% CI shown per program.

Finding The top-6 programs (Stanford, Wharton, HBS, MIT Sloan, Columbia, Kellogg) each exceed 0.57, a tier boundary roughly 3x above Tuck (0.280, rank 9). The gap between rank 8 (Booth, 0.442) and rank 9 (Tuck, 0.280) is the sharpest discontinuity in the distribution and likely reflects a M7 citation threshold in the training corpus.
M₁ᴺᴳᴹᴼ (TDK, training memory) M₁ᴬᴳˢᴼ (LSS, live search) ΔM₁ = LSS − TDK
1 Stanford GSB 0.627 +0.027
2 Wharton 0.625 −0.025
3 Harvard Business School 0.603 −0.002
4 MIT Sloan 0.597 +0.000
5 Columbia Business School 0.592 +0.002
6 Kellogg 0.573 −0.025
7 Haas 0.473 −0.022
8 Booth 0.442 +0.000
9 Tuck 0.280 −0.045
10 Yale SOM 0.275 +0.007
11 Stern 0.245 −0.005
12 Ross 0.188 −0.045
13 Fuqua 0.177 −0.013
14 Darden 0.122 +0.000
15 Anderson 0.117 −0.013
16 Tepper 0.087 −0.010
17 Marshall 0.073 −0.005
18 Foster 0.062 +0.002
19 McCombs 0.060 −0.028
20 Johnson (Cornell) 0.048 −0.022
21 Kenan-Flagler 0.045 +0.012
22 Mendoza / Olin 0.020 +0.000
24 Carey 0.015 +0.002
25 Smith (Maryland) 0.013 +0.002

Bars scale to the corpus maximum (0.627). Gray = TDK (training memory) · Blue = LSS (live search). ΔM₁ chip: green = search-amplified, gray = stable. Source: MBA-AEPI v1.9.0-FROZEN, September 2026.

Divergence analysis

Training data vs. live search: the ΔM₁ story

ΔM₁ = M₁ᴬᴳˢᴼ − M₁ᴺᴳᴹᴼ. Positive = live search amplifies the program beyond model memory. Option D dual-gate: |ΔM₁| ≥ 0.05 AND non-overlapping Wilson 95% CIs required for reclassification.

TDK Mode
Training Data Knowledge
No web search, model memory only
600 prompts collected
600Prompt Pairs
Each Run
Twice
CSPS Protocol
LSS Mode
Live Search Surface
Web search enabled, retrieval active
600 prompts collected
ΔM₁ Output
Amplified · Stable · Suppressed
All 25: Stable at Wave 1
Key finding All 25 programs are classified Stable. This is the expected result for a single engine at n=600: the Option D dual-gate (|ΔM₁| ≥ 0.05 AND non-overlapping CIs) is deliberately conservative. The notable exception is MIT Sloan's Factual stratum: M₁ᴬᴳˢᴼ = 0.200 vs. M₁ᴺᴳᴹᴼ = 0.067, a 3x within-stratum search amplification. Cross-engine expansion in Waves 2 to 5 will increase resolution and enable the first divergence classifications.
Program M₁ᴬᴳˢᴼ LSS M₁ᴺᴳᴹᴼ TDK ΔM₁ |ΔM₁| Class
Stanford GSB0.6270.600+0.0270.027Stable
Wharton0.6250.650−0.0250.025Stable
Harvard Business School0.6030.605−0.0020.002Stable
MIT Sloan0.5970.597+0.0000.000Stable
Columbia Business School0.5920.590+0.0020.002Stable
Kellogg0.5730.598−0.0250.025Stable
Haas0.4730.495−0.0220.022Stable
Booth0.4420.442+0.0000.000Stable
Tuck0.2800.325−0.0450.045Stable
Yale SOM0.2750.268+0.0070.007Stable
Stern0.2450.250−0.0050.005Stable
Ross0.1880.233−0.0450.045Stable
Fuqua0.1770.190−0.0130.013Stable
Darden0.1220.122+0.0000.000Stable
Anderson0.1170.130−0.0130.013Stable

Top 15 programs shown. Full table of 25 in the Technical Appendix.

Act 2 How AI decides
Cross-stratum robustness

Engine Consistency Metric (ECM)

ECM = 1 minus CV(stratum M₁ᴬᴳˢᴼ scores). Range [0,1]. Higher = more consistent mention rate across Generic, Constraints, Persona, and Factual strata. Low ECM = prompt-fragile visibility.

Finding MIT Sloan leads ECM at 0.63. It is the only program that maintains meaningful visibility in Factual-stratum prompts (0.200 LSS vs. 0.050 baseline for most programs). McCombs and Kenan-Flagler have ECM ≥ 0.51 despite low headline scores, indicating consistent cross-stratum presence. Foster (ECM 0.00) and Johnson (ECM 0.06) are effectively single-stratum programs, nearly invisible outside their primary prompt type.
Program ECM Generic Constraints Persona Factual Profile
MIT Sloan0.63 0.6210.6940.6000.200Stratum-robust
McCombs0.52 0.0880.0560.0170.050Consistent low
Kenan-Flagler0.51 0.0500.0440.0170.083Consistent low
Haas0.49 0.5170.5330.5080.050Factual drop-off
Wharton0.48 0.6580.7170.7080.050Factual drop-off
Booth0.48 0.5500.4440.4170.050Generic-led
Harvard Business School0.47 0.6000.7280.7000.050Factual drop-off
Columbia Business School0.47 0.6830.6890.5330.050Factual drop-off
Anderson0.45 0.1170.1830.0500.050Constraints-led
Kellogg0.46 0.5460.6670.7500.050Persona-led
Stanford GSB0.46 0.6000.7390.8000.050Persona-led
Marshall0.42 0.1000.0890.0080.050Generic-led
Fuqua0.41 0.1040.2280.3080.050Persona-led
Darden0.39 0.1540.0440.2080.050Generic/Persona
Tuck0.38 0.4330.1560.2750.050Generic-led
What this means A program with high M₁ but low ECM is competitively fragile: its AI visibility concentrates in one prompt type and collapses in others. Stanford GSB's Persona score of 0.800 (highest in the corpus) masks a Factual score of 0.050. Programs seeking to close the gap should invest in exactly the strata where they underperform.
Sentiment analysis

Valence: ChatGPT as recommender

Valence scored via keyword-pattern matching in ±200-character context windows around each entity mention. Positive = laudatory framing or list-position signal. Negative = critical or cautionary language.

0%
Negative mentions, corpus-wide
Zero negative framing detected across all 1,200 responses and 25 programs
40.8%
Positive mention ratio
59.2% neutral. "Ranked," "leading," and "prestigious" language dominate
0.644
Foster: highest valence ratio
47 positive of 73 total mentions. Small n amplifies ratio.
0.640
Harvard Business School valence
464 positive of 725 total, the largest absolute positive-mention count
Finding ChatGPT operates as a pure recommender at Wave 1. It names programs, ranks them, and describes their attributes without critical or cautionary framing. This is structurally significant: the AI surface is not a neutral information channel, it is an advocacy channel. Programs that appear in AI responses are implicitly endorsed. Programs that do not appear are implicitly excluded.
Act 3 What this means for strategy
Framework

AEO vs. SEO: why AI visibility is a distinct discipline

Search Engine Optimization (SEO) targets indexing signals. AI Engine Optimization (AEO) operates on a fundamentally different mechanism: parametric memory compressed during training, sampled during generation.

The core distinction A program ranking #1 in Google for "top MBA programs" may still be structurally invisible to GPT-4o if its web presence did not generate sufficient co-citation mass in the training corpus. MBA-AEPI quantifies this gap. Columbia Business School (AI rank 5) and Kellogg (rank 6) surpass programs with comparable traditional prestige in specific prompt contexts. The AI surface rewards programs described frequently, positively, and in association with specific outcome signals in publicly accessible, high-authority web text.
Strategic interpretation

What your Stable classification means

Stable is not "safe." It is a statement that live-search and model-memory mention rates are statistically consistent within current measurement resolution. The strategic implication depends entirely on where in the index a program sits.

ClassificationProfileStrategic implication
High M₁ + StableStanford, Wharton, HBS, MIT Sloan, Columbia, KelloggDeeply embedded in both TDK and LSS. Maintenance requires sustained output in AI-indexed channels.
Mid M₁ + Low ECMYale SOM, Tuck, BoothPrompt-fragile: strong in generic queries, weak in constrained or factual ones. Targeted content strategy addresses this.
Low M₁ + StableMendoza, Olin, Carey, SmithInvisible in both modes. Must build AI citation mass from near zero. Longest investment horizon.
Act 4 What comes next
Multi-engine roadmap

Wave 2: Perplexity, coming September 10, 2026

Wave 2 launches September 10, 2026, replicating the full MBA-AEPI v1.9.0-FROZEN protocol against Perplexity AI and enabling the first cross-engine comparative index. All five platforms are on track to be live by September 30, 2026.

Wave 2 additions Perplexity's native inline citation architecture provides structured source metadata (domain, URL, date) unavailable in ChatGPT's non-citation responses. Wave 2 adds: cross-engine ΔM₁ (Spearman rank correlation), LLM-based valence scoring (replacing the keyword-window proxy), and a full citation-authority analysis by domain type.
Engine Wave Status Architecture Est. delivery
ChatGPT (GPT-4o)Wave 1✓ CompleteTDK + LSS (web search toggle)September 2026
Perplexity AIWave 2PlannedNative citation architectureSeptember 10, 2026
Claude (Anthropic)Wave 3PlannedTDK + web retrievalBy September 30, 2026
Microsoft CopilotWave 4PlannedBing-grounded retrievalBy September 30, 2026
Google GeminiWave 5PlannedGoogle Search groundedBy September 30, 2026
Infrastructure How this is scored
Protocol

Methodology: CSPS v1.9.0-FROZEN

All MBA-AEPI waves use the Concurrent Synchronised Prompt Sequence (CSPS) protocol: TDK and LSS captures for each prompt pair, executed within the same session and 24-hour window.

600
Prompt pairs
240 Generic · 180 Constraints · 120 Persona · 60 Factual
1,200
Total captures
600 TDK + 600 LSS · 100% HTTP 200 · 0 failures
25
Target programs
All 25 detected ≥1 response · 0 zero-mention programs
Core scoring formulas

M₁ᴬᴳˢᴼ = mentions_LSS / 600: headline LSS mention rate

M₁ᴺᴳᴹᴼ = mentions_TDK / 600: TDK mention rate (model memory)

ΔM₁ = M₁ᴬᴳˢᴼ − M₁ᴺᴳᴹᴼ: divergence

ECM = max(0, 1 − σ/μ): consistency across 4 strata (σ=SD, μ=mean)

Wilson CI: (p̂ + z²/2n ± z·√(p̂(1−p̂)/n + z²/4n²)) / (1 + z²/n): α=0.05, z=1.96

ΔM₁ Option D: Search-Amplified if |ΔM₁| ≥ 0.05 AND CIs non-overlapping AND ΔM₁ > 0

Technical appendix

Data governance & complete entity table

Full 25-program table with disambiguation notes. All figures drawn directly from MBA-AEPI v1.9.0-FROZEN production run, September 2, 2026.

Rk Program M₁ᴬᴳˢᴼ 95% CI M₁ᴺᴳᴹᴼ ΔM₁ ECM Valence Mentions
1Stanford GSB0.627[0.587, 0.664]0.600+0.0270.460.62736
2Wharton0.625[0.586, 0.663]0.650−0.0250.480.59765
3Harvard Business School0.603[0.564, 0.642]0.605−0.0020.470.64725
4MIT Sloan0.597[0.557, 0.635]0.597+0.0000.630.33716
5Columbia Business School0.592[0.552, 0.630]0.590+0.0020.470.37709
6Kellogg0.573[0.533, 0.612]0.598−0.0250.460.27703
7Haas0.473[0.434, 0.513]0.495−0.0220.490.33581
8Booth0.442[0.402, 0.482]0.442+0.0000.480.32530
9Tuck0.280[0.246, 0.317]0.325−0.0450.380.32363
10Yale SOM0.275[0.241, 0.312]0.268+0.0070.220.46326
11Stern0.245[0.212, 0.281]0.250−0.0050.430.36297
12Ross0.188[0.159, 0.222]0.233−0.0450.320.19253
13Fuqua0.177[0.148, 0.209]0.190−0.0130.410.25220
14Darden0.122[0.098, 0.150]0.122+0.0000.390.32146
15Anderson0.117[0.093, 0.145]0.130−0.0130.450.24148
16Tepper0.087[0.067, 0.112]0.097−0.0100.120.32110
17Marshall0.073[0.055, 0.097]0.078−0.0050.420.3391
18Foster0.062[0.045, 0.084]0.060+0.0020.000.6473
19McCombs0.060[0.044, 0.082]0.088−0.0280.520.3789
20Johnson (Cornell)0.048[0.034, 0.069]0.070−0.0220.060.3971
21Kenan-Flagler0.045[0.031, 0.065]0.033+0.0120.510.2847
22Mendoza0.020[0.011, 0.035]0.018+0.0020.000.2623
22Olin (WashU)0.020[0.011, 0.035]0.020+0.0000.000.0824
24Carey (Johns Hopkins)0.015[0.008, 0.028]0.013+0.0020.000.5317
25Smith (Maryland)0.013[0.007, 0.026]0.012+0.0020.000.3315

Source: MBA-AEPI Wave 1 production run · ChatGPT GPT-4o · September 2, 2026 · v1.9.0-FROZEN · 1,200 captures · 100% completion · 0 anomalies

Common questions about MBA-AEPI

What is MBA-AEPI?

MBA-AEPI (MBA AI Engine Perception Index) measures how often ChatGPT names each of 25 MBA programs across two modes: what the model already knows from training, and what it says once it can search the live web. Wave 1 covers ChatGPT (GPT-4o) only, with 1,200 responses across 600 dual-mode prompt pairs.

How is MBA-AEPI different from Arcalea's M7 AEO Index?

They are separate instruments. The M7 Index tracks M7 business schools across five AI platforms with a five-signal composite score and a memory-versus-retrieval quadrant. MBA-AEPI Wave 1 tracks a broader 25-program cohort on ChatGPT only, using a training-memory-versus-live-search comparison and a different scoring model built around Wilson confidence intervals. Neither replaces the other; MBA-AEPI is a newer, narrower baseline that will expand across engines in later waves.

What does a Stable classification mean?

Stable means a program's live-search and training-memory mention rates are not reliably different once measurement error is accounted for. It is the default result at Wave 1: with a single engine and n=600, the Option D rule (a confidence-interval gap of at least 0.05, with non-overlapping intervals) is deliberately conservative, so most programs land here.

Why does Wave 1 use ChatGPT only?

MBA-AEPI is built as a longitudinal instrument, adding one engine at a time so each wave's methodology is fully validated before the next is layered on. ChatGPT (GPT-4o) is the Wave 1 baseline. Perplexity launches Wave 2 on September 10, 2026, followed by Claude, Copilot, and Gemini, all targeted to be live by September 30, 2026.

How is AI visibility different from a traditional MBA ranking?

A traditional ranking (US News, Financial Times) scores programs on inputs like selectivity, salary outcomes, and survey results. MBA-AEPI scores something else: how often and how consistently ChatGPT actually names a program when someone asks it for guidance. A program can rank well in traditional guides and still be nearly invisible in AI answers, or the reverse, since AI visibility depends on how much citable, co-referenced content about a program exists in the model's training data and on the live web.

What is the Engine Consistency Metric (ECM)?

ECM measures how evenly a program is mentioned across four prompt types: Generic, Constraints, Persona, and Factual. A program with a high headline score but low ECM is winning in one context and nearly absent in others, which makes its AI visibility fragile rather than durable.