MBA-AEPI (MBA AI Engine Perception Index) measures how often GPT-4o, accessed through the OpenAI API, names each of 25 MBA programs in answers to a fixed 600-prompt registry under specified configurations. This ChatGPT release covers three configurations: a no-tool baseline and an optional-search condition (September 2, 2026) and a required-search condition (September 5 UTC), 1,800 responses in total. It measures AI visibility, how often each program is named in generated answers under those configurations, reported alongside established quality rankings.
The AI-visibility index for MBA programs.
Two search configurations, reported separately: how often GPT-4o names each program with web search off (a no-tool baseline) and with web search enabled, first optional and then required. Two labeled rankings, paired statistical tests, and the limits stated alongside every number. ChatGPT (GPT-4o API) only.
Executive Summary
Everything this index measures, in ninety seconds. Every figure is drawn from the panels below, so the summary and the evidence cannot disagree.
What the statistics mean
Every technical term used on this page, defined for a non-specialist reader. Hover any underlined term elsewhere on the page for the short version. For a reader's guide to evaluating any AEO report, see How to Read an AEO Analysis.
| Term | What it means here |
|---|---|
| Three configurations | No-tool (NGMO, also called TDK): GPT-4o answers with web search switched off. Optional search (AGSO): search is available and the model decides whether to use it. Required search (AGSO_FORCED): the model must search before answering. Each has its own mention rate; the two search conditions produce the two rankings on this page, always labeled. |
| M₁, the mention rate | The share of 600 prompts on which a program was named at least once. 0.627 means Stanford GSB appeared in 62.7% of optional-search answers. It counts presence, not prominence, praise, or accuracy. |
| Wilson 95% confidence interval | The range within which the true mention rate plausibly sits, given that 600 prompts is a sample. Tighter is more precise. It covers sampling error only, not changes in the model or the prompt set over time, and it is not a test between two programs. |
| ΔM₁ | A program's mention rate with search minus its rate without. Negative means the program was named less often when search was on. It is a measured difference between two configurations, labeled by condition. |
| Paired exact McNemar test | Because the same 600 prompts ran in both configurations, each prompt is compared with itself. The test asks whether prompts that flipped from named to not-named outnumber those that flipped the other way by more than chance would produce. Only the prompts that flipped (the discordant pairs) carry information. |
| Benjamini-Hochberg correction and the q value | Testing 25 programs at once inflates false alarms. The Benjamini-Hochberg step adjusts each result into a q value; a program is labeled Search-Suppressed or Search-Amplified only when q is 0.05 or below. Values near the line (McCombs and Ross at q = 0.0505) are shown to four decimals for that reason. |
| Stable · Search-Suppressed · Search-Amplified | The three statistical labels. Stable means the difference did not reach the q ≤ 0.05 threshold in this sample; equivalence is a separate question the test leaves open. Suppressed and Amplified indicate direction of a significant difference, not its practical size. |
| Statistical vs. practical materiality | Significance says a difference is unlikely to be chance; materiality asks whether it is large enough to matter. This release reports effect sizes alongside labels; the separate paired-interval materiality assessment described in Amendment v1.9.1 has not been performed. |
| Option D (legacy) | The earlier rule of thumb: non-overlapping intervals and a gap of at least five percentage points. Kept as a reference diagnostic only; the dotted ±0.05 lines on the ΔM₁ charts mark it. Labels come from the McNemar/BH test. |
| Spearman rho | How closely the AI visibility order matches the U.S. News order, from −1 to 1; 1.0 would be identical ordering. 0.973 is a strong association. It shows correspondence, not cause. |
| Rank-on-rank OLS and R² | A straight-line fit of AI rank on U.S. News rank. R² (0.944) is the share of the variation in AI rank that the fit accounts for; the remaining 5.6% is what the line leaves unexplained. The score-space fit on raw mention rates is retained for comparison but models a different quantity and predicts impossible negative rates from U.S. News rank 22 downward. |
| Rank premium and deficit | How many places better (premium) or worse (deficit) a program's AI rank is than the fitted line predicts from its U.S. News rank. Columbia sits 3.48 places better than predicted; Johnson 3.90 places worse. Positions relative to a fitted line within this cohort, on one collection date. |
| Competition ranking | Programs with exactly equal mention rates share a rank and the next rank number is skipped: Mendoza and Olin are both 22, so Carey is 24 and no program is 23. |
| Prompt strata | The 600 prompts fall into four designed types: Generic (240), Constraints (180), Persona (120), and Factual (60). All 60 Factual prompts name schools, so the full index is a mixed discovery/lookup index; the 540-prompt sensitivity analysis repeats the analysis without them. |
| Cross-Stratum Consistency (CSC) | How evenly a program's mention rate is spread across the four prompt types, from 0 (highly uneven) to 1 (perfectly even). Because the Factual stratum names schools with unequal coverage, CSC describes unevenness across the designed prompt mix, not fragility. |
| Valence proxy | A keyword scan of the text around the first mention of each program in each answer, tagging laudatory words as positive and cautionary words as negative. Pooled across the 1,200 original responses. A rough indicator of framing, not a validated sentiment model. |
| Search invocation vs. citation | Invocation means the model actually ran a web search; a citation is a source link in the answer. Optional search did not log invocation, so citation presence (83 of 600) is a proxy. Required search set an invocation flag on all 600 responses; 583 carried citations. |
| Model alias | gpt-4o is a name that can point to different underlying model versions over time; no snapshot identifier was recorded. Comparisons across the two collection dates therefore include possible model drift. |
| CSPS | Collect Simultaneously, Publish Sequentially: the design rule that paired conditions are collected together. The no-tool and optional-search conditions were; the required-search condition (September 5) is a documented exception against the reused September 2 baseline. |
Find your program
Pull every figure for one program into a single view instead of hunting across the sections below.
The M₁ᴬᴳˢᴼ headline index (optional search)
M₁ᴬᴳˢᴼ = proportion of the 600 responses collected September 2, 2026 with web search enabled but not required, in which each program is named at least once. Wilson 95% CI shown per program. Search invocation was not logged in this condition; 83 of 600 responses (13.8%) contained citations. The 600 prompts are not all entity-free: 540 are discovery-style (Generic, Constraints, Persona) and 60 Factual prompts explicitly name schools, so this is a mixed discovery/lookup index (see the registry note under Methodology). The forced-search ranking is reported separately below.
The chart below plots that same story with every program's Wilson 95% CI attached. The intervals show binomial uncertainty around each individual mention rate; they are not paired between-program tests and do not account for prompt selection or dependence.
The breakdown below adds what the chart above can't show: how each program's optional-search rate compares with the no-tool baseline collected in the same September 2 session, program by program, with the exact gap between the two. Both are generated outputs under specified configurations; the baseline is not a direct measurement of training-data contents.
Bars scale to the corpus maximum (0.627). Gray = TDK (no-tool baseline) · Blue = LSS (optional search). ΔM₁ chip: blue = positive (optional search above the no-tool baseline), gray = zero or negative. Original optional-search condition only. Source: mba_aepi_index_scores_wave1_chatgpt_v2.csv, September 2, 2026.
AI mention rank vs. US News rank
How closely does the original optional-search ranking track the US News MBA rank recorded for the cohort? This is a descriptive comparison of two orderings, not a causal model. These figures have not been recomputed for the forced-search condition.
The regression itself, all 25 programs plotted against US News rank with the OLS fit line:
| Metric | Rank-space (AI rank ~ US News rank) | Score-space (M₁ᴬᴳˢᴼ ~ US News rank) |
|---|---|---|
| Spearman rho | 0.973 | −0.977 |
| Pearson r | n/a | −0.941 |
| R² | 0.944 (adj. 0.941) | 0.886 |
| Unexplained | 5.6% | 11.4% |
| p-value | < 0.001 | < 0.001 |
Rank-space is the primary read (see callout above); score-space is retained for comparison with the original analysis. Both fits use the original optional-search condition only. No correlation statistics have been computed for the forced-search condition in this release. Source: correlation_summary_v2.json.
Rank premium = predicted AI rank minus actual AI rank, from the rank-space regression above. Positive means a program surfaces at a better AI rank than its US News prestige would predict; negative means worse. All 25 programs below, sorted from largest premium to largest deficit. This replaces the earlier score-space residual leaderboard, which produced infeasible negative predicted mention rates at the bottom of the distribution. Full regression methodology in the Technical Validation Annex.
No-tool baseline vs. optional search: the ΔM₁ story
ΔM₁ = M₁ᴬᴳˢᴼ − M₁ᴺᴳᴹᴼ, a measured contrast between two generated outputs for the same prompt. Positive means more mentions in the search-enabled configuration; it does not by itself establish a causal retrieval effect. The primary statistical label is a paired exact McNemar test with Benjamini-Hochberg correction across 25 programs. The older Option D CI-overlap rule is retained only as a legacy diagnostic and is shown on the chart for reference.
Web search off; a research construct, not a training-data readout
600 responses, Sept 2
Each Run
Twice
Web search enabled; invocation not logged, 83/600 cited
600 responses, Sept 2
Original: 25 Stable · Forced: 13 / 3 / 9
Original optional-search ΔM₁ for all 25 programs, sorted. Every point sits inside ±0.05; the largest are Tuck and Ross at −0.045 and Stanford at +0.027. The ±0.05 lines mark the legacy Option D threshold for reference only. Statistical labels come from the paired McNemar/BH test, not from this gate, and no paired-interval materiality assessment is supplied in this release.
| Program | M₁ᴬᴳˢᴼ optional search | M₁ᴺᴳᴹᴼ no-tool | ΔM₁ | |ΔM₁| | BH class |
|---|---|---|---|---|---|
| Stanford GSB | 0.627 | 0.600 | +0.027 | 0.027 | Stable |
| Wharton | 0.625 | 0.650 | −0.025 | 0.025 | Stable |
| Harvard Business School | 0.603 | 0.605 | −0.002 | 0.002 | Stable |
| MIT Sloan | 0.558 | 0.573 | −0.015 | 0.015 | Stable |
| Columbia Business School | 0.592 | 0.590 | +0.002 | 0.002 | Stable |
| Kellogg | 0.573 | 0.598 | −0.025 | 0.025 | Stable |
| Haas | 0.473 | 0.495 | −0.022 | 0.022 | Stable |
| Booth | 0.442 | 0.442 | +0.000 | 0.000 | Stable |
| Tuck | 0.280 | 0.325 | −0.045 | 0.045 | Stable |
| Yale SOM | 0.275 | 0.268 | +0.007 | 0.007 | Stable |
| Stern | 0.245 | 0.250 | −0.005 | 0.005 | Stable |
| Ross | 0.188 | 0.233 | −0.045 | 0.045 | Stable |
| Fuqua | 0.177 | 0.190 | −0.013 | 0.013 | Stable |
| Darden | 0.122 | 0.122 | +0.000 | 0.000 | Stable |
| Anderson | 0.117 | 0.130 | −0.013 | 0.013 | Stable |
Top 15 by original optional-search M₁ shown; all 25 are Stable in this condition. Full table in the Technical Appendix. Forced-search results follow in the required-search section.
Cross-Stratum Consistency (CSC)
CSC = 1 minus CV(stratum M₁ᴬᴳˢᴼ scores). Range [0,1]. Higher = more even mention rate across the four designed strata (Generic, Constraints, Persona, Factual); lower = more uneven. Because the Factual stratum consists of named-school lookups with unequal coverage, CSC describes unevenness across the strata as designed, not established fragility or broad discovery consistency.
| Program | CSC | Generic | Constraints | Persona | Factual | Profile |
|---|---|---|---|---|---|---|
| McCombs | 0.52 | 0.088 | 0.056 | 0.017 | 0.050 | Consistent low |
| Kenan-Flagler | 0.43 | 0.050 | 0.011 | 0.017 | 0.050 | Consistent low |
| Haas | 0.49 | 0.517 | 0.533 | 0.508 | 0.050 | Factual drop-off |
| Booth | 0.48 | 0.550 | 0.444 | 0.417 | 0.050 | Generic-led |
| MIT Sloan | 0.48 | 0.579 | 0.683 | 0.583 | 0.050 | Constraints-led |
| Wharton | 0.48 | 0.658 | 0.717 | 0.708 | 0.050 | Factual drop-off |
| Harvard Business School | 0.47 | 0.600 | 0.728 | 0.700 | 0.050 | Factual drop-off |
| Columbia Business School | 0.47 | 0.683 | 0.689 | 0.533 | 0.050 | Factual drop-off |
| Anderson | 0.45 | 0.117 | 0.183 | 0.050 | 0.050 | Constraints-led |
| Kellogg | 0.46 | 0.546 | 0.667 | 0.750 | 0.050 | Persona-led |
| Stanford GSB | 0.46 | 0.600 | 0.739 | 0.800 | 0.050 | Persona-led |
| Marshall | 0.42 | 0.100 | 0.089 | 0.008 | 0.050 | Generic-led |
| Fuqua | 0.41 | 0.104 | 0.228 | 0.308 | 0.050 | Persona-led |
| Darden | 0.39 | 0.154 | 0.044 | 0.208 | 0.050 | Generic/Persona |
| Tuck | 0.38 | 0.433 | 0.156 | 0.275 | 0.050 | Generic-led |
Valence: keyword-proxy framing (pooled no-tool and optional-search responses)
Valence scored via keyword-pattern matching in ±200-character context windows around the first matched occurrence for each program in each response. Positive = laudatory framing or list-position signal. Negative = critical or cautionary language.
AEO vs. SEO: why AI visibility is a distinct discipline
Search Engine Optimization (SEO) concerns how a page ranks in a results list. AI Engine Optimization (AEO) concerns how an entity is represented in generated answers. They measure different outcomes while sharing much of the same technical and content foundation; retrieval-based AEO still depends on search infrastructure.
Reading a program's profile: investigation paths, not diagnoses
A classification label describes a measured relationship between two configurations on the dates collected. Stable means the difference did not reach the BH significance threshold in this sample; equivalence is a separate question the test leaves open. The patterns below are places to look next. None is an intervention this study tested.
| Observed pattern | Programs (original optional search) | What to investigate first |
|---|---|---|
| High M₁, Stable originally, suppressed under forced search | Stanford, Wharton, HBS, Columbia, Kellogg, MIT Sloan | Replicate the drop with both conditions collected close together and the model snapshot pinned. A large gap between runs can reflect configuration, timing, sampling variation, or retrieval; it does not by itself establish dependence on one information source. |
| Mid M₁, low CSC | Yale SOM, Tuck, Booth | Review the actual constraint and factual questions and answers. The program may not meet those constraints, its attributes may be unclear, or the detector may miss mentions. Content coverage is one hypothesis to test, not a conclusion from the gap alone. |
| Low M₁ in both configurations | Mendoza, Olin, Carey, Smith | First validate entity detection and prompt relevance, then check whether accurate, accessible information about the program exists. A low mention rate locates an absence; the cause is the next thing to investigate. |
| Significant but small increase under forced search | Carey (+0.030), Foster (+0.028), Smith (+0.012) | Inspect the returned citations and the answers' factual claims. Each increase is BH-significant but under five percentage points and has not passed a separate materiality assessment. |
What changed when search was required
The required-search condition reran the same 600 prompt IDs with the search tool set to required, so that web search was invoked on every call, and logged invocation per response. It reused the original September 2 no-tool baseline; no new baseline was collected alongside it. The comparison therefore describes observed differences between two runs three days apart under an unpinned gpt-4o model alias. The recovered collection scripts also differ in one request setting: the original defaults to a 1,500-token output cap (overridable on the command line; the historical invocation is not archived), while the forced-search script sets no explicit cap. It does not fully isolate retrieval from timing, model-version, or request-setting differences.
Citation presence by prompt type, original optional search vs. forced search. The Persona stratum, which had no citations in the original condition, produced cited answers in 115 of 120 forced-search responses. That confirms these prompts can yield retrieval; it does not retrospectively establish whether search ran in the original uncited answers.
| Stratum | Optional search, with citations | Forced search, with citations | Forced search, invocation logged |
|---|---|---|---|
| Generic | 14 / 240 (5.8%) | 235 / 240 (97.9%) | 240 / 240 |
| Constraints | 19 / 180 (10.6%) | 174 / 180 (96.7%) | 180 / 180 |
| Persona | 0 / 120 (0.0%) | 115 / 120 (95.8%) | 120 / 120 |
| Factual | 50 / 60 (83.3%) | 59 / 60 (98.3%) | 60 / 60 |
| All strata | 83 / 600 (13.8%) | 583 / 600 (97.2%) | 600 / 600 |
Forced-search ΔM₁ for all 25 programs, sorted. Compare the scale with the original chart above: the largest original difference was −0.045, while here 12 programs fall past −0.05 (the legacy Option D threshold, shown for reference) and a thirteenth, Anderson at −0.045, is BH-significant without crossing it. Labels come from the paired McNemar/BH test.
The forced-search ranking. Ranked on M₁ᴬᴳˢᴼ under required search. "Rank, optional" is the same program's rank in the original optional-search condition; "Change" is optional rank minus forced rank. Columbia and Booth tie at 0.360 and share rank 7.
| Rk | Program | M₁ᴺᴳᴹᴼ no-tool | M₁ forced search | ΔM₁ | Rank, optional | Change | BH class |
|---|---|---|---|---|---|---|---|
| 1 | Wharton | 0.650 | 0.447 | −0.203 | 2 | +1 | Search-Suppressed |
| 2 | Kellogg | 0.598 | 0.408 | −0.190 | 5 | +3 | Search-Suppressed |
| 3 | Harvard Business School | 0.605 | 0.387 | −0.218 | 3 | 0 | Search-Suppressed |
| 4 | MIT Sloan | 0.573 | 0.385 | −0.188 | 6 | +2 | Search-Suppressed |
| 5 | Stanford GSB | 0.600 | 0.383 | −0.217 | 1 | -4 | Search-Suppressed |
| 6 | Haas | 0.495 | 0.382 | −0.113 | 7 | +1 | Search-Suppressed |
| 7 | Booth | 0.442 | 0.360 | −0.082 | 8 | +1 | Search-Suppressed |
| 7 | Columbia Business School | 0.590 | 0.360 | −0.230 | 4 | -3 | Search-Suppressed |
| 9 | Stern | 0.250 | 0.235 | −0.015 | 11 | +2 | Stable |
| 10 | Yale SOM | 0.268 | 0.188 | −0.080 | 10 | 0 | Search-Suppressed |
Top 10 by forced-search M₁ shown. All 25, with confidence intervals and the legacy Option D label alongside the primary BH label, are in the Technical Appendix.
Did the findings depend on the named-school factual prompts?
Because all 60 Factual prompts name schools, we excluded them symmetrically from both arms of each comparison and repeated the analysis on the remaining 540 prompt IDs (240 Generic, 180 Constraints, 120 Persona; effective weights 44.4% / 33.3% / 22.2%, no reweighting). Regex detection is unchanged; Wilson intervals use n = 540; exact McNemar tests and BH correction are rerun within each comparison. This is a sensitivity analysis, not a new pre-specified index, and the two analyses share responses, so it is not an independent replication.
| Rk | Program (forced-search top six) | M₁, 600 prompts | M₁, 540 prompts | Δ vs. no-tool, 600 | Δ vs. no-tool, 540 | BH class (both) |
|---|---|---|---|---|---|---|
| 1 | Wharton | 0.447 | 0.491 | -20.33 pp | -22.59 pp | Search-Suppressed |
| 2 | Kellogg | 0.408 | 0.448 | -19.00 pp | -21.11 pp | Search-Suppressed |
| 3 | Harvard Business School | 0.387 | 0.424 | -21.83 pp | -24.26 pp | Search-Suppressed |
| 4 | MIT Sloan | 0.385 | 0.422 | -18.83 pp | -20.93 pp | Search-Suppressed |
| 5 | Stanford GSB | 0.383 | 0.420 | -21.67 pp | -24.07 pp | Search-Suppressed |
| 6 | Haas | 0.382 | 0.419 | -11.33 pp | -12.59 pp | Search-Suppressed |
Largest negative difference on 540 prompts remains Columbia (-25.56 pp); the three amplified programs remain Carey (+3.33 pp), Foster (+2.96 pp), and Smith (+1.30 pp). Only Foster and Smith have changed BH q values (0.0305 and 0.0260), both still below 0.05; all optional-search p/q values are identical because the excluded Factual observations contributed no discordant pairs. Source: sensitivity/comparison_600_vs_540.csv and sensitivity/summary.json.
Methodology: v1.9.0-FROZEN with Amendment v1.9.1
The no-tool and optional-search conditions each collected one response per prompt for 600 prompts in one session on September 2, 2026 (CSPS protocol). The required-search condition is a documented exception: its 600 responses are timestamped September 5 UTC and are compared against the reused September 2 baseline, so prompt pairing controls the question asked but not the collection date. Statistical labels follow Amendment v1.9.1: paired exact McNemar with Benjamini-Hochberg correction across 25 programs; the legacy Option D rule is retained separately.
M₁ᴬᴳˢᴼ = mentions_search / 600: search-enabled mention rate, labeled by condition
M₁ᴺᴳᴹᴼ = mentions_no-tool / 600: no-tool baseline mention rate
ΔM₁ = M₁ᴬᴳˢᴼ − M₁ᴺᴳᴹᴼ: configuration contrast
CSC = max(0, 1 − σ/μ): consistency across 4 strata (σ=SD, μ=mean)
Wilson CI: (p̂ + z²/2n ± z·√(p̂(1−p̂)/n + z²/4n²)) / (1 + z²/n): α=0.05, z=1.96
Primary label: exact McNemar on discordant pairs (b = no-tool only, c = search only) → BH q across 25; Amplified if q ≤ 0.05 and ΔM₁ > 0, Suppressed if q ≤ 0.05 and ΔM₁ < 0, else Stable
Legacy Option D (diagnostic only): |ΔM₁| ≥ 0.05 AND non-overlapping Wilson CIs
Data governance & complete entity tables
Table A: original optional-search condition (September 2, 2026). Ranked on M₁ᴬᴳˢᴼ with optional search. Source: mba_aepi_index_scores_wave1_chatgpt_v2.csv. Table B, the forced-search condition, follows.
| Rk | Program | M₁ᴬᴳˢᴼ | 95% CI | M₁ᴺᴳᴹᴼ | ΔM₁ | CSC | Valence | Mentions |
|---|---|---|---|---|---|---|---|---|
| 1 | Stanford GSB | 0.627 | [0.587, 0.664] | 0.600 | +0.027 | 0.46 | 0.62 | 736 |
| 2 | Wharton | 0.625 | [0.586, 0.663] | 0.650 | −0.025 | 0.48 | 0.59 | 765 |
| 3 | Harvard Business School | 0.603 | [0.564, 0.642] | 0.605 | −0.002 | 0.47 | 0.64 | 725 |
| 4 | Columbia Business School | 0.592 | [0.552, 0.630] | 0.590 | +0.002 | 0.47 | 0.37 | 709 |
| 5 | Kellogg | 0.573 | [0.533, 0.612] | 0.598 | −0.025 | 0.46 | 0.27 | 703 |
| 6 | MIT Sloan | 0.558 | [0.518, 0.598] | 0.573 | −0.015 | 0.48 | 0.33 | 679 |
| 7 | Haas | 0.473 | [0.434, 0.513] | 0.495 | −0.022 | 0.49 | 0.33 | 581 |
| 8 | Booth | 0.442 | [0.402, 0.482] | 0.442 | +0.000 | 0.48 | 0.32 | 530 |
| 9 | Tuck | 0.280 | [0.246, 0.317] | 0.325 | −0.045 | 0.38 | 0.32 | 363 |
| 10 | Yale SOM | 0.275 | [0.241, 0.312] | 0.268 | +0.007 | 0.22 | 0.46 | 326 |
| 11 | Stern | 0.245 | [0.212, 0.281] | 0.250 | −0.005 | 0.43 | 0.36 | 297 |
| 12 | Ross | 0.188 | [0.159, 0.222] | 0.233 | −0.045 | 0.32 | 0.19 | 253 |
| 13 | Fuqua | 0.177 | [0.148, 0.209] | 0.190 | −0.013 | 0.41 | 0.25 | 220 |
| 14 | Darden | 0.122 | [0.098, 0.150] | 0.122 | +0.000 | 0.39 | 0.32 | 146 |
| 15 | Anderson | 0.117 | [0.093, 0.145] | 0.130 | −0.013 | 0.45 | 0.24 | 148 |
| 16 | Tepper | 0.087 | [0.067, 0.112] | 0.097 | −0.010 | 0.12 | 0.32 | 110 |
| 17 | Marshall | 0.073 | [0.055, 0.097] | 0.078 | −0.005 | 0.42 | 0.33 | 91 |
| 18 | Foster | 0.062 | [0.045, 0.084] | 0.060 | +0.002 | 0.00 | 0.64 | 73 |
| 19 | McCombs | 0.060 | [0.044, 0.082] | 0.088 | −0.028 | 0.52 | 0.37 | 89 |
| 20 | Johnson (Cornell) | 0.048 | [0.034, 0.069] | 0.070 | −0.022 | 0.06 | 0.39 | 71 |
| 21 | Kenan-Flagler | 0.032 | [0.020, 0.049] | 0.028 | +0.003 | 0.43 | 0.28 | 36 |
| 22 | Mendoza | 0.020 | [0.011, 0.035] | 0.018 | +0.002 | 0.00 | 0.26 | 23 |
| 22 | Olin (WashU) | 0.020 | [0.011, 0.035] | 0.020 | +0.000 | 0.00 | 0.08 | 24 |
| 24 | Carey (Johns Hopkins) | 0.015 | [0.008, 0.028] | 0.013 | +0.002 | 0.00 | 0.53 | 17 |
| 25 | Smith (Maryland) | 0.013 | [0.007, 0.026] | 0.012 | +0.002 | 0.00 | 0.33 | 15 |
Rank 22 repeats and rank 23 is skipped, not a typo: Mendoza and Olin are exactly tied at M₁ᴬᴳˢᴼ = 0.020 (12 of 600 responses each) and share rank 22, so the next distinct value (Carey, 0.015) takes rank 24. Standard competition-ranking convention.
Source, Table A: MBA-AEPI · ChatGPT GPT-4o (alias, snapshot not recorded) · September 2, 2026 · 1,200 captures · 100% completion · response length 207 to 5,964 characters (median 1,588.5). All 25 programs are Stable under the BH rule in this condition.
Table B: forced-search condition
Ranked on M₁ᴬᴳˢᴼ with the search tool set to required, collected September 5, 2026 UTC, against the reused September 2 no-tool baseline. "BH class" is the primary label (paired exact McNemar, Benjamini-Hochberg across 25, α 0.05). "Legacy Option D" is the older CI-overlap/five-point diagnostic, shown for continuity only. Source: mba_aepi_index_scores_wave1b_chatgpt_v3.csv.
| Rk | Program | M₁ forced | 95% CI | M₁ᴺᴳᴹᴼ | ΔM₁ | Rk, optional | BH class | Legacy Option D |
|---|---|---|---|---|---|---|---|---|
| 1 | Wharton | 0.447 | [0.407, 0.487] | 0.650 | −0.203 | 2 | Search-Suppressed | Search-Suppressed |
| 2 | Kellogg | 0.408 | [0.370, 0.448] | 0.598 | −0.190 | 5 | Search-Suppressed | Search-Suppressed |
| 3 | Harvard Business School | 0.387 | [0.349, 0.426] | 0.605 | −0.218 | 3 | Search-Suppressed | Search-Suppressed |
| 4 | MIT Sloan | 0.385 | [0.347, 0.425] | 0.573 | −0.188 | 6 | Search-Suppressed | Search-Suppressed |
| 5 | Stanford GSB | 0.383 | [0.345, 0.423] | 0.600 | −0.217 | 1 | Search-Suppressed | Search-Suppressed |
| 6 | Haas | 0.382 | [0.344, 0.421] | 0.495 | −0.113 | 7 | Search-Suppressed | Search-Suppressed |
| 7 | Booth | 0.360 | [0.323, 0.399] | 0.442 | −0.082 | 8 | Search-Suppressed | Search-Suppressed |
| 7 | Columbia Business School | 0.360 | [0.323, 0.399] | 0.590 | −0.230 | 4 | Search-Suppressed | Search-Suppressed |
| 9 | Stern | 0.235 | [0.203, 0.271] | 0.250 | −0.015 | 11 | Stable | Stable |
| 10 | Yale SOM | 0.188 | [0.159, 0.222] | 0.268 | −0.080 | 10 | Search-Suppressed | Search-Suppressed |
| 11 | Ross | 0.148 | [0.122, 0.179] | 0.233 | −0.085 | 12 | Search-Suppressed | Search-Suppressed |
| 12 | Tuck | 0.133 | [0.108, 0.163] | 0.325 | −0.192 | 9 | Search-Suppressed | Search-Suppressed |
| 13 | Darden | 0.118 | [0.095, 0.147] | 0.122 | −0.003 | 14 | Stable | Stable |
| 14 | Fuqua | 0.115 | [0.092, 0.143] | 0.190 | −0.075 | 13 | Search-Suppressed | Search-Suppressed |
| 15 | Johnson | 0.093 | [0.073, 0.119] | 0.070 | +0.023 | 20 | Stable | Stable |
| 16 | Foster | 0.088 | [0.068, 0.114] | 0.060 | +0.028 | 18 | Search-Amplified | Stable |
| 17 | Anderson | 0.085 | [0.065, 0.110] | 0.130 | −0.045 | 15 | Search-Suppressed | Stable |
| 18 | Marshall | 0.082 | [0.062, 0.106] | 0.078 | +0.003 | 17 | Stable | Stable |
| 19 | McCombs | 0.080 | [0.061, 0.104] | 0.088 | −0.008 | 19 | Stable | Stable |
| 20 | Tepper | 0.077 | [0.058, 0.101] | 0.097 | −0.020 | 16 | Stable | Stable |
| 21 | Carey | 0.043 | [0.030, 0.063] | 0.013 | +0.030 | 24 | Search-Amplified | Stable |
| 22 | Kenan-Flagler | 0.023 | [0.014, 0.039] | 0.028 | −0.005 | 21 | Stable | Stable |
| 22 | Smith | 0.023 | [0.014, 0.039] | 0.012 | +0.012 | 25 | Search-Amplified | Stable |
| 24 | Olin | 0.017 | [0.009, 0.030] | 0.020 | −0.003 | 22 | Stable | Stable |
| 25 | Mendoza | 0.005 | [0.002, 0.015] | 0.018 | −0.013 | 22 | Stable | Stable |
Ranks 7 and 22 repeat: Columbia and Booth tie at 0.360; Smith and Kenan-Flagler tie at 0.023. Competition-ranking convention. Legacy Option D yields 12 suppressed / 13 Stable; it is not the headline classification.
Source, Table B: MBA-AEPI · ChatGPT GPT-4o · September 5, 2026, 01:14 to 02:16 UTC · 600 captures · invocation flag 600/600 · citations in 583/600 (97.2%), mean 6.79 annotations per response. Provenance: both raw ChatGPT collections (1,200 original responses timestamped Sept 2, 16:59 to 18:49 UTC; 600 forced-search responses), the 600-prompt registry with SHA-256 hashes, and both collection scripts are retained in the research repository and fingerprinted in chatgpt_source_manifest.json; they are not included in the downloads on this page, which are limited to the methodology and companion documents. The read-only verifier reconstructs every mention count, both paired-count matrices, all p/q values, and all 600 registry hashes from the raw text. Open items it leaves for future work are detector accuracy against human labels, the exact historical command-line settings or SDK versions, and capture-time hash validation, which the recovered scripts did not perform.
Common questions about MBA-AEPI
No, not all of them. The 540 Generic, Constraints, and Persona prompts passed an automated scan for target-school names, but all 60 Factual prompts explicitly name schools (three variants each for 20 schools, 19 of them tracked). Earlier descriptions calling every prompt entity-free were incorrect. The aggregate scores are therefore a mixed discovery/lookup index. Scores were not changed; a separate 540-prompt robustness check is reported on the page.
In these saved responses, no. Excluding all 60 Factual prompts left the original optional-search ranking unchanged for all 25 programs, kept all 25 Stable, and left the forced-search top six and all 13 / 3 / 9 BH classifications unchanged, with three lower-table rank moves (Tepper, McCombs, Kenan-Flagler). Rates rise mechanically because the denominator shrinks. This is a post-hoc robustness check, not an independent replication.
There are two, and they differ. In the original optional-search condition Stanford GSB led at 0.627, followed by Wharton 0.625, Harvard 0.603, Columbia 0.592, Kellogg 0.573, and MIT Sloan 0.558. In the forced-search condition Wharton led at 0.447, followed by Kellogg 0.408, Harvard 0.387, MIT Sloan 0.385, and Stanford 0.383. Always label the condition; neither is simply the ChatGPT rank.
Stable means the program's mention-rate difference between two configurations did not reach significance under a paired exact McNemar test with Benjamini-Hochberg correction across 25 programs at alpha 0.05. Whether the configurations are equivalent is a separate question that this test leaves open. In the original optional-search condition all 25 programs were Stable; under forced search, 13 were Search-Suppressed, 3 Search-Amplified, and 9 Stable. Statistical significance is separate from practical materiality, which this release does not assess.
For the original optional-search condition, invocation was not logged; citations appeared in 83 of 600 responses (13.8%), and citation presence is a proxy, not a measured invocation rate. The required-search condition set an invocation flag on all 600 responses; that flag is set by either a web_search_call item or a citation annotation and is not a retained raw tool-call trace. 583 responses (97.2%) contained citations. Zero citations do not imply a zero difference in mention rates: Stanford's Persona rate was 0.800 with optional search versus 0.617 without, despite no citations.
Both raw ChatGPT collections, the 600-prompt registry with hashes, both collection scripts, the scoring code, and all outputs are retained in the research repository; they are not included in the downloads on this page, which are limited to the methodology and companion documents. A read-only verifier reconstructs every mention count, both paired comparisons, all p/q values, and all registry hashes from the raw text without API calls. Open items it leaves for future work are detector accuracy against human labels, the exact historical command-line settings or SDK versions, and future model outputs.
The release scope is ChatGPT under no-tool, optional-search, and required-search conditions. Results for other engines and any cross-engine comparison are outside this release and are not reported here. The next step for this instrument is to replicate the forced-search comparison with closely timed conditions, explicit invocation logging, a pinned model snapshot, and human-validated entity detection.
A traditional ranking (US News, Financial Times) scores programs on inputs like selectivity, salary outcomes, and survey results. MBA-AEPI scores something else: how often GPT-4o names a program in answers to a fixed prompt set under a specified configuration. In the original optional-search condition that ranking tracked the recorded US News rank closely (Spearman rho 0.973), while individual programs still sat above or below the fitted line. The study does not inspect training data or identify which sources shaped an answer, so it does not explain why a program appears more or less often.
CSC (Cross-Stratum Consistency) measures how evenly a program's mention rate holds across the four query strata: Generic, Constraints, Persona, and Factual. A high CSC means the program is named consistently across query types; a low CSC means its mention rates are uneven across the four designed strata. CSC = 1 minus the ratio of standard deviation to mean across the four stratum M1 values, clipped to [0, 1]. CSC values on this page are from the original optional-search condition; note that the Factual stratum is a named-school lookup stratum.