MBA-AEPI (MBA AI Engine Perception Index) measures how often ChatGPT names each of 25 MBA programs across two modes: what the model already knows from training, and what it says once it can search the live web. Wave 1 covers ChatGPT (GPT-4o) only, with 1,200 responses across 600 dual-mode prompt pairs.
The complete AI-visibility index for MBA programs.
Two modes of AI visibility: what ChatGPT already knows from training, and what it says once it can search live. Measured separately, scored honestly, and compared in one view. Wave 1: ChatGPT GPT-4o.
Executive Summary
Everything this index measures, in ninety seconds. Every figure is drawn from the panels below, so the summary and the evidence cannot disagree.
The M₁ᴬᴳˢᴼ headline index
M₁ᴬᴳˢᴼ = proportion of ChatGPT Live Search Surface responses in which each program is mentioned. n=600 prompts. Wilson 95% CI shown per program.
Bars scale to the corpus maximum (0.627). Gray = TDK (training memory) · Blue = LSS (live search). ΔM₁ chip: green = search-amplified, gray = stable. Source: MBA-AEPI v1.9.0-FROZEN, September 2026.
Training data vs. live search: the ΔM₁ story
ΔM₁ = M₁ᴬᴳˢᴼ − M₁ᴺᴳᴹᴼ. Positive = live search amplifies the program beyond model memory. Option D dual-gate: |ΔM₁| ≥ 0.05 AND non-overlapping Wilson 95% CIs required for reclassification.
No web search, model memory only
600 prompts collected
Each Run
Twice
Web search enabled, retrieval active
600 prompts collected
All 25: Stable at Wave 1
| Program | M₁ᴬᴳˢᴼ LSS | M₁ᴺᴳᴹᴼ TDK | ΔM₁ | |ΔM₁| | Class |
|---|---|---|---|---|---|
| Stanford GSB | 0.627 | 0.600 | +0.027 | 0.027 | Stable |
| Wharton | 0.625 | 0.650 | −0.025 | 0.025 | Stable |
| Harvard Business School | 0.603 | 0.605 | −0.002 | 0.002 | Stable |
| MIT Sloan | 0.597 | 0.597 | +0.000 | 0.000 | Stable |
| Columbia Business School | 0.592 | 0.590 | +0.002 | 0.002 | Stable |
| Kellogg | 0.573 | 0.598 | −0.025 | 0.025 | Stable |
| Haas | 0.473 | 0.495 | −0.022 | 0.022 | Stable |
| Booth | 0.442 | 0.442 | +0.000 | 0.000 | Stable |
| Tuck | 0.280 | 0.325 | −0.045 | 0.045 | Stable |
| Yale SOM | 0.275 | 0.268 | +0.007 | 0.007 | Stable |
| Stern | 0.245 | 0.250 | −0.005 | 0.005 | Stable |
| Ross | 0.188 | 0.233 | −0.045 | 0.045 | Stable |
| Fuqua | 0.177 | 0.190 | −0.013 | 0.013 | Stable |
| Darden | 0.122 | 0.122 | +0.000 | 0.000 | Stable |
| Anderson | 0.117 | 0.130 | −0.013 | 0.013 | Stable |
Top 15 programs shown. Full table of 25 in the Technical Appendix.
Engine Consistency Metric (ECM)
ECM = 1 minus CV(stratum M₁ᴬᴳˢᴼ scores). Range [0,1]. Higher = more consistent mention rate across Generic, Constraints, Persona, and Factual strata. Low ECM = prompt-fragile visibility.
| Program | ECM | Generic | Constraints | Persona | Factual | Profile |
|---|---|---|---|---|---|---|
| MIT Sloan | 0.63 | 0.621 | 0.694 | 0.600 | 0.200 | Stratum-robust |
| McCombs | 0.52 | 0.088 | 0.056 | 0.017 | 0.050 | Consistent low |
| Kenan-Flagler | 0.51 | 0.050 | 0.044 | 0.017 | 0.083 | Consistent low |
| Haas | 0.49 | 0.517 | 0.533 | 0.508 | 0.050 | Factual drop-off |
| Wharton | 0.48 | 0.658 | 0.717 | 0.708 | 0.050 | Factual drop-off |
| Booth | 0.48 | 0.550 | 0.444 | 0.417 | 0.050 | Generic-led |
| Harvard Business School | 0.47 | 0.600 | 0.728 | 0.700 | 0.050 | Factual drop-off |
| Columbia Business School | 0.47 | 0.683 | 0.689 | 0.533 | 0.050 | Factual drop-off |
| Anderson | 0.45 | 0.117 | 0.183 | 0.050 | 0.050 | Constraints-led |
| Kellogg | 0.46 | 0.546 | 0.667 | 0.750 | 0.050 | Persona-led |
| Stanford GSB | 0.46 | 0.600 | 0.739 | 0.800 | 0.050 | Persona-led |
| Marshall | 0.42 | 0.100 | 0.089 | 0.008 | 0.050 | Generic-led |
| Fuqua | 0.41 | 0.104 | 0.228 | 0.308 | 0.050 | Persona-led |
| Darden | 0.39 | 0.154 | 0.044 | 0.208 | 0.050 | Generic/Persona |
| Tuck | 0.38 | 0.433 | 0.156 | 0.275 | 0.050 | Generic-led |
Valence: ChatGPT as recommender
Valence scored via keyword-pattern matching in ±200-character context windows around each entity mention. Positive = laudatory framing or list-position signal. Negative = critical or cautionary language.
AEO vs. SEO: why AI visibility is a distinct discipline
Search Engine Optimization (SEO) targets indexing signals. AI Engine Optimization (AEO) operates on a fundamentally different mechanism: parametric memory compressed during training, sampled during generation.
What your Stable classification means
Stable is not "safe." It is a statement that live-search and model-memory mention rates are statistically consistent within current measurement resolution. The strategic implication depends entirely on where in the index a program sits.
| Classification | Profile | Strategic implication |
|---|---|---|
| High M₁ + Stable | Stanford, Wharton, HBS, MIT Sloan, Columbia, Kellogg | Deeply embedded in both TDK and LSS. Maintenance requires sustained output in AI-indexed channels. |
| Mid M₁ + Low ECM | Yale SOM, Tuck, Booth | Prompt-fragile: strong in generic queries, weak in constrained or factual ones. Targeted content strategy addresses this. |
| Low M₁ + Stable | Mendoza, Olin, Carey, Smith | Invisible in both modes. Must build AI citation mass from near zero. Longest investment horizon. |
Wave 2: Perplexity, coming September 10, 2026
Wave 2 launches September 10, 2026, replicating the full MBA-AEPI v1.9.0-FROZEN protocol against Perplexity AI and enabling the first cross-engine comparative index. All five platforms are on track to be live by September 30, 2026.
| Engine | Wave | Status | Architecture | Est. delivery |
|---|---|---|---|---|
| ChatGPT (GPT-4o) | Wave 1 | ✓ Complete | TDK + LSS (web search toggle) | September 2026 |
| Perplexity AI | Wave 2 | Planned | Native citation architecture | September 10, 2026 |
| Claude (Anthropic) | Wave 3 | Planned | TDK + web retrieval | By September 30, 2026 |
| Microsoft Copilot | Wave 4 | Planned | Bing-grounded retrieval | By September 30, 2026 |
| Google Gemini | Wave 5 | Planned | Google Search grounded | By September 30, 2026 |
Methodology: CSPS v1.9.0-FROZEN
All MBA-AEPI waves use the Concurrent Synchronised Prompt Sequence (CSPS) protocol: TDK and LSS captures for each prompt pair, executed within the same session and 24-hour window.
M₁ᴬᴳˢᴼ = mentions_LSS / 600: headline LSS mention rate
M₁ᴺᴳᴹᴼ = mentions_TDK / 600: TDK mention rate (model memory)
ΔM₁ = M₁ᴬᴳˢᴼ − M₁ᴺᴳᴹᴼ: divergence
ECM = max(0, 1 − σ/μ): consistency across 4 strata (σ=SD, μ=mean)
Wilson CI: (p̂ + z²/2n ± z·√(p̂(1−p̂)/n + z²/4n²)) / (1 + z²/n): α=0.05, z=1.96
ΔM₁ Option D: Search-Amplified if |ΔM₁| ≥ 0.05 AND CIs non-overlapping AND ΔM₁ > 0
Data governance & complete entity table
Full 25-program table with disambiguation notes. All figures drawn directly from MBA-AEPI v1.9.0-FROZEN production run, September 2, 2026.
| Rk | Program | M₁ᴬᴳˢᴼ | 95% CI | M₁ᴺᴳᴹᴼ | ΔM₁ | ECM | Valence | Mentions |
|---|---|---|---|---|---|---|---|---|
| 1 | Stanford GSB | 0.627 | [0.587, 0.664] | 0.600 | +0.027 | 0.46 | 0.62 | 736 |
| 2 | Wharton | 0.625 | [0.586, 0.663] | 0.650 | −0.025 | 0.48 | 0.59 | 765 |
| 3 | Harvard Business School | 0.603 | [0.564, 0.642] | 0.605 | −0.002 | 0.47 | 0.64 | 725 |
| 4 | MIT Sloan | 0.597 | [0.557, 0.635] | 0.597 | +0.000 | 0.63 | 0.33 | 716 |
| 5 | Columbia Business School | 0.592 | [0.552, 0.630] | 0.590 | +0.002 | 0.47 | 0.37 | 709 |
| 6 | Kellogg | 0.573 | [0.533, 0.612] | 0.598 | −0.025 | 0.46 | 0.27 | 703 |
| 7 | Haas | 0.473 | [0.434, 0.513] | 0.495 | −0.022 | 0.49 | 0.33 | 581 |
| 8 | Booth | 0.442 | [0.402, 0.482] | 0.442 | +0.000 | 0.48 | 0.32 | 530 |
| 9 | Tuck | 0.280 | [0.246, 0.317] | 0.325 | −0.045 | 0.38 | 0.32 | 363 |
| 10 | Yale SOM | 0.275 | [0.241, 0.312] | 0.268 | +0.007 | 0.22 | 0.46 | 326 |
| 11 | Stern | 0.245 | [0.212, 0.281] | 0.250 | −0.005 | 0.43 | 0.36 | 297 |
| 12 | Ross | 0.188 | [0.159, 0.222] | 0.233 | −0.045 | 0.32 | 0.19 | 253 |
| 13 | Fuqua | 0.177 | [0.148, 0.209] | 0.190 | −0.013 | 0.41 | 0.25 | 220 |
| 14 | Darden | 0.122 | [0.098, 0.150] | 0.122 | +0.000 | 0.39 | 0.32 | 146 |
| 15 | Anderson | 0.117 | [0.093, 0.145] | 0.130 | −0.013 | 0.45 | 0.24 | 148 |
| 16 | Tepper | 0.087 | [0.067, 0.112] | 0.097 | −0.010 | 0.12 | 0.32 | 110 |
| 17 | Marshall | 0.073 | [0.055, 0.097] | 0.078 | −0.005 | 0.42 | 0.33 | 91 |
| 18 | Foster | 0.062 | [0.045, 0.084] | 0.060 | +0.002 | 0.00 | 0.64 | 73 |
| 19 | McCombs | 0.060 | [0.044, 0.082] | 0.088 | −0.028 | 0.52 | 0.37 | 89 |
| 20 | Johnson (Cornell) | 0.048 | [0.034, 0.069] | 0.070 | −0.022 | 0.06 | 0.39 | 71 |
| 21 | Kenan-Flagler | 0.045 | [0.031, 0.065] | 0.033 | +0.012 | 0.51 | 0.28 | 47 |
| 22 | Mendoza | 0.020 | [0.011, 0.035] | 0.018 | +0.002 | 0.00 | 0.26 | 23 |
| 22 | Olin (WashU) | 0.020 | [0.011, 0.035] | 0.020 | +0.000 | 0.00 | 0.08 | 24 |
| 24 | Carey (Johns Hopkins) | 0.015 | [0.008, 0.028] | 0.013 | +0.002 | 0.00 | 0.53 | 17 |
| 25 | Smith (Maryland) | 0.013 | [0.007, 0.026] | 0.012 | +0.002 | 0.00 | 0.33 | 15 |
Source: MBA-AEPI Wave 1 production run · ChatGPT GPT-4o · September 2, 2026 · v1.9.0-FROZEN · 1,200 captures · 100% completion · 0 anomalies
Common questions about MBA-AEPI
They are separate instruments. The M7 Index tracks M7 business schools across five AI platforms with a five-signal composite score and a memory-versus-retrieval quadrant. MBA-AEPI Wave 1 tracks a broader 25-program cohort on ChatGPT only, using a training-memory-versus-live-search comparison and a different scoring model built around Wilson confidence intervals. Neither replaces the other; MBA-AEPI is a newer, narrower baseline that will expand across engines in later waves.
Stable means a program's live-search and training-memory mention rates are not reliably different once measurement error is accounted for. It is the default result at Wave 1: with a single engine and n=600, the Option D rule (a confidence-interval gap of at least 0.05, with non-overlapping intervals) is deliberately conservative, so most programs land here.
MBA-AEPI is built as a longitudinal instrument, adding one engine at a time so each wave's methodology is fully validated before the next is layered on. ChatGPT (GPT-4o) is the Wave 1 baseline. Perplexity launches Wave 2 on September 10, 2026, followed by Claude, Copilot, and Gemini, all targeted to be live by September 30, 2026.
A traditional ranking (US News, Financial Times) scores programs on inputs like selectivity, salary outcomes, and survey results. MBA-AEPI scores something else: how often and how consistently ChatGPT actually names a program when someone asks it for guidance. A program can rank well in traditional guides and still be nearly invisible in AI answers, or the reverse, since AI visibility depends on how much citable, co-referenced content about a program exists in the model's training data and on the live web.
ECM measures how evenly a program is mentioned across four prompt types: Generic, Constraints, Persona, and Factual. A program with a high headline score but low ECM is winning in one context and nearly absent in others, which makes its AI visibility fragile rather than durable.