MBA-AEPI · First release · September 2026 · ChatGPT (GPT-4o API)

The AI-visibility index for MBA programs.

Two search configurations, reported separately: how often GPT-4o names each program with web search off (a no-tool baseline) and with web search enabled, first optional and then required. Two labeled rankings, paired statistical tests, and the limits stated alongside every number. ChatGPT (GPT-4o API) only.

25 MBA programs 600 prompt IDs · 1,800 responses No-tool · optional search · required search Two rankings, both labeled
New instrument MBA-AEPI is a companion instrument to the AEO Index: M7 Business Schools, with a different methodology, a ChatGPT-only scope, and a broader 25-program cohort. This release covers ChatGPT under no-tool, optional-search, and required-search conditions; it does not reanalyze the M7 dataset or rely on its findings, and it does not replace that index.
25
Programs tracked
Stanford 0.627
Original optional-search leader
Wharton 0.447
Forced-search leader
13 · 3 · 9
Forced search: suppressed · amplified · Stable
1,800
Captures across two collections
600/600
Invocation logged (required search)
0%
Negative-keyword matches
ChatGPT release Three configurations, one first release
The short version

Executive Summary

Everything this index measures, in ninety seconds. Every figure is drawn from the panels below, so the summary and the evidence cannot disagree.

Finding This release reports two distinct ChatGPT rankings and labels the condition every time. Original optional search (Sept 2): Stanford GSB 0.627, Wharton 0.625, and Harvard Business School 0.603 lead and their Wilson 95% CIs overlap; below Booth (rank 8, 0.442) mention rates drop sharply: 17 programs rank below Booth and 10 of 25 are under 0.10. Forced search (Sept 5 UTC): Wharton 0.447, Kellogg 0.408, Harvard 0.387, MIT Sloan 0.385, Stanford 0.383 lead, and 13 programs were named significantly less often than in the reused no-tool baseline. Both rankings measure the same thing: how often GPT-4o, accessed through the OpenAI API, names each program under a specified search configuration. They report AI visibility alongside the U.S. News quality ranking, and the two are compared directly in the prestige section.
What this means M₁ is a mention rate over a mixed prompt set (540 discovery-style prompts plus 60 named-school factual lookups), not a quality ranking or a measurement of what the model was trained on. In the original optional-search condition it tracks the recorded US News rank closely (Spearman rho 0.973) while still differing program by program: Columbia (0.592) is named more often than Kellogg (0.573) despite comparable traditional standing, and McCombs' CSC of 0.52 is the most even spread across prompt strata despite a low headline score. The required-search condition produced a different ranking and lower rates for most programs; the two conditions are reported separately throughout and neither is simply "the ChatGPT rank."
Key terms, in plain language

What the statistics mean

Every technical term used on this page, defined for a non-specialist reader. Hover any underlined term elsewhere on the page for the short version. For a reader's guide to evaluating any AEO report, see How to Read an AEO Analysis.

TermWhat it means here
Three configurationsNo-tool (NGMO, also called TDK): GPT-4o answers with web search switched off. Optional search (AGSO): search is available and the model decides whether to use it. Required search (AGSO_FORCED): the model must search before answering. Each has its own mention rate; the two search conditions produce the two rankings on this page, always labeled.
M₁, the mention rateThe share of 600 prompts on which a program was named at least once. 0.627 means Stanford GSB appeared in 62.7% of optional-search answers. It counts presence, not prominence, praise, or accuracy.
Wilson 95% confidence intervalThe range within which the true mention rate plausibly sits, given that 600 prompts is a sample. Tighter is more precise. It covers sampling error only, not changes in the model or the prompt set over time, and it is not a test between two programs.
ΔM₁A program's mention rate with search minus its rate without. Negative means the program was named less often when search was on. It is a measured difference between two configurations, labeled by condition.
Paired exact McNemar testBecause the same 600 prompts ran in both configurations, each prompt is compared with itself. The test asks whether prompts that flipped from named to not-named outnumber those that flipped the other way by more than chance would produce. Only the prompts that flipped (the discordant pairs) carry information.
Benjamini-Hochberg correction and the q valueTesting 25 programs at once inflates false alarms. The Benjamini-Hochberg step adjusts each result into a q value; a program is labeled Search-Suppressed or Search-Amplified only when q is 0.05 or below. Values near the line (McCombs and Ross at q = 0.0505) are shown to four decimals for that reason.
Stable · Search-Suppressed · Search-AmplifiedThe three statistical labels. Stable means the difference did not reach the q ≤ 0.05 threshold in this sample; equivalence is a separate question the test leaves open. Suppressed and Amplified indicate direction of a significant difference, not its practical size.
Statistical vs. practical materialitySignificance says a difference is unlikely to be chance; materiality asks whether it is large enough to matter. This release reports effect sizes alongside labels; the separate paired-interval materiality assessment described in Amendment v1.9.1 has not been performed.
Option D (legacy)The earlier rule of thumb: non-overlapping intervals and a gap of at least five percentage points. Kept as a reference diagnostic only; the dotted ±0.05 lines on the ΔM₁ charts mark it. Labels come from the McNemar/BH test.
Spearman rhoHow closely the AI visibility order matches the U.S. News order, from −1 to 1; 1.0 would be identical ordering. 0.973 is a strong association. It shows correspondence, not cause.
Rank-on-rank OLS and R²A straight-line fit of AI rank on U.S. News rank. R² (0.944) is the share of the variation in AI rank that the fit accounts for; the remaining 5.6% is what the line leaves unexplained. The score-space fit on raw mention rates is retained for comparison but models a different quantity and predicts impossible negative rates from U.S. News rank 22 downward.
Rank premium and deficitHow many places better (premium) or worse (deficit) a program's AI rank is than the fitted line predicts from its U.S. News rank. Columbia sits 3.48 places better than predicted; Johnson 3.90 places worse. Positions relative to a fitted line within this cohort, on one collection date.
Competition rankingPrograms with exactly equal mention rates share a rank and the next rank number is skipped: Mendoza and Olin are both 22, so Carey is 24 and no program is 23.
Prompt strataThe 600 prompts fall into four designed types: Generic (240), Constraints (180), Persona (120), and Factual (60). All 60 Factual prompts name schools, so the full index is a mixed discovery/lookup index; the 540-prompt sensitivity analysis repeats the analysis without them.
Cross-Stratum Consistency (CSC)How evenly a program's mention rate is spread across the four prompt types, from 0 (highly uneven) to 1 (perfectly even). Because the Factual stratum names schools with unequal coverage, CSC describes unevenness across the designed prompt mix, not fragility.
Valence proxyA keyword scan of the text around the first mention of each program in each answer, tagging laudatory words as positive and cautionary words as negative. Pooled across the 1,200 original responses. A rough indicator of framing, not a validated sentiment model.
Search invocation vs. citationInvocation means the model actually ran a web search; a citation is a source link in the answer. Optional search did not log invocation, so citation presence (83 of 600) is a proxy. Required search set an invocation flag on all 600 responses; 583 carried citations.
Model aliasgpt-4o is a name that can point to different underlying model versions over time; no snapshot identifier was recorded. Comparisons across the two collection dates therefore include possible model drift.
CSPSCollect Simultaneously, Publish Sequentially: the design rule that paired conditions are collected together. The no-tool and optional-search conditions were; the required-search condition (September 5) is a documented exception against the reused September 2 baseline.
Quick lookup

Find your program

Pull every figure for one program into a single view instead of hunting across the sections below.

25 programs, alphabetical
Act 1 Where programs stand
Primary metric · original optional-search condition

The M₁ᴬᴳˢᴼ headline index (optional search)

M₁ᴬᴳˢᴼ = proportion of the 600 responses collected September 2, 2026 with web search enabled but not required, in which each program is named at least once. Wilson 95% CI shown per program. Search invocation was not logged in this condition; 83 of 600 responses (13.8%) contained citations. The 600 prompts are not all entity-free: 540 are discovery-style (Generic, Constraints, Persona) and 60 Factual prompts explicitly name schools, so this is a mixed discovery/lookup index (see the registry note under Methodology). The forced-search ranking is reported separately below.

What M₁ measures M₁ measures presence: whether a program was named at least once in a response. Separate dimensions, reserved for later instruments, include recommendation strength, position or prominence within a response, citation support, attribute association, share of answer, or influence on enrollment behavior.
Finding The top six (Stanford 0.627, Wharton 0.625, HBS 0.603, Columbia 0.592, Kellogg 0.573, MIT Sloan 0.558) span 0.558 to 0.627, roughly twice Tuck (0.280, rank 9). The gap between rank 8 (Booth, 0.442) and rank 9 (Tuck, 0.280) is the sharpest discontinuity in the distribution and is the only tier break with clearly non-overlapping intervals. This study does not inspect training data, so it does not identify why the break sits where it does. MIT Sloan's score was revised from 0.597 (rank 4) to 0.558 (rank 6) after an entity-detection correction; see Methodology §6.1.

The chart below plots that same story with every program's Wilson 95% CI attached. The intervals show binomial uncertainty around each individual mention rate; they are not paired between-program tests and do not account for prompt selection or dependence.

Color marks three descriptive bands in the distribution, not a statistical test on any pair. Two adjacent pairs have non-overlapping intervals (MIT Sloan and Haas and Booth and Tuck); the mid-tier/long-tail break is a visual grouping only, since Anderson's and Tepper's CIs overlap. #1 is ringed in amber.
Top tier (rank 1 to 8) Mid tier (9 to 15) Long tail (16 to 25)

The breakdown below adds what the chart above can't show: how each program's optional-search rate compares with the no-tool baseline collected in the same September 2 session, program by program, with the exact gap between the two. Both are generated outputs under specified configurations; the baseline is not a direct measurement of training-data contents.

M₁ᴺᴳᴹᴼ (TDK, no-tool baseline) M₁ᴬᴳˢᴼ (LSS, optional search) ΔM₁ = LSS − TDK
1 Stanford GSB 0.627 +0.027
2 Wharton 0.625 −0.025
3 Harvard Business School 0.603 −0.002
4 Columbia Business School 0.592 +0.002
5 Kellogg 0.573 −0.025
6 MIT Sloan 0.558 −0.015
7 Haas 0.473 −0.022
8 Booth 0.442 +0.000
9 Tuck 0.280 −0.045
10 Yale SOM 0.275 +0.007
11 Stern 0.245 −0.005
12 Ross 0.188 −0.045
13 Fuqua 0.177 −0.013
14 Darden 0.122 +0.000
15 Anderson 0.117 −0.013
16 Tepper 0.087 −0.010
17 Marshall 0.073 −0.005
18 Foster 0.062 +0.002
19 McCombs 0.060 −0.028
20 Johnson (Cornell) 0.048 −0.022
21 Kenan-Flagler 0.032 +0.003
22 Mendoza 0.020 +0.002
22 Olin 0.020 +0.000
24 Carey 0.015 +0.002
25 Smith (Maryland) 0.013 +0.002

Bars scale to the corpus maximum (0.627). Gray = TDK (no-tool baseline) · Blue = LSS (optional search). ΔM₁ chip: blue = positive (optional search above the no-tool baseline), gray = zero or negative. Original optional-search condition only. Source: mba_aepi_index_scores_wave1_chatgpt_v2.csv, September 2, 2026.

What this means This is the metric a program cannot see on any traditional dashboard. Website analytics measure visits; this index measures program mentions in generated answers. M₁ᴬᴳˢᴼ measures whether a program appears in answers to this fixed mix of discovery-style questions and named-school lookups. Low visibility identifies a pattern to investigate; applicant exposure, website visits, and enrollment effects are separate questions this study did not measure.
Cross-check · original optional-search condition

AI mention rank vs. US News rank

How closely does the original optional-search ranking track the US News MBA rank recorded for the cohort? This is a descriptive comparison of two orderings, not a causal model. These figures have not been recomputed for the forced-search condition.

Finding The original optional-search AI rank is strongly associated with the recorded US News rank. The primary specification is rank-on-rank OLS (AI rank vs. US News rank, both bounded 1 to 25): Spearman rho = 0.973, R² = 0.944 (adj. 0.941), p < 0.001, leaving 5.6% of rank variance unexplained by that fit. A score-space OLS on the raw M₁ proportions (Pearson r = −0.941, R² = 0.886, 11.4% unexplained) is retained for comparison, but it models a different outcome, predicts negative mention rates from U.S. News rank 22 downward (Mendoza, Smith, Olin, Carey), and its S-curve residual pattern is a property of the functional form, so its unexplained share is not interchangeable with the rank-space figure. Neither number is a measured "independent AEO effect"; both are descriptive fit statistics for one condition on one date.

The regression itself, all 25 programs plotted against US News rank with the OLS fit line:

Program (M₁ vs. rank) Score-space OLS fit (R²=0.886)
MetricRank-space (AI rank ~ US News rank)Score-space (M₁ᴬᴳˢᴼ ~ US News rank)
Spearman rho0.973−0.977
Pearson rn/a−0.941
0.944 (adj. 0.941)0.886
Unexplained5.6%11.4%
p-value< 0.001< 0.001

Rank-space is the primary read (see callout above); score-space is retained for comparison with the original analysis. Both fits use the original optional-search condition only. No correlation statistics have been computed for the forced-search condition in this release. Source: correlation_summary_v2.json.

Rank premium = predicted AI rank minus actual AI rank, from the rank-space regression above. Positive means a program surfaces at a better AI rank than its US News prestige would predict; negative means worse. All 25 programs below, sorted from largest premium to largest deficit. This replaces the earlier score-space residual leaderboard, which produced infeasible negative predicted mention rates at the bottom of the distribution. Full regression methodology in the Technical Validation Annex.

AI-rank premium AI-rank deficit
What this means Columbia Business School is the standout: ranked 7th by US News, its original optional-search AI rank of 4 is 3.48 places better than the fit predicts, the largest premium in the cohort. Tepper, Foster, Olin, and Stanford GSB also appear at a better AI rank than predicted. On the deficit side, Johnson shows the largest gap (3.90 places worse than predicted), followed by Smith, Booth, and McCombs. MIT Sloan is a modest under-index (1.40 places worse than predicted) once entity detection is corrected; an earlier version of this analysis, using an inflated MIT Sloan count and the score-space fit, had reported it as over-indexed. These are positions relative to a fitted line within this cohort on one collection date. They are not evidence about program merit, about the causes of visibility, or about how a program would rank under forced search, where Columbia fell to rank 7 (tied) at 0.360.
Divergence analysis · original optional-search condition

No-tool baseline vs. optional search: the ΔM₁ story

ΔM₁ = M₁ᴬᴳˢᴼ − M₁ᴺᴳᴹᴼ, a measured contrast between two generated outputs for the same prompt. Positive means more mentions in the search-enabled configuration; it does not by itself establish a causal retrieval effect. The primary statistical label is a paired exact McNemar test with Benjamini-Hochberg correction across 25 programs. The older Option D CI-overlap rule is retained only as a legacy diagnostic and is shown on the chart for reference.

TDK Mode
No-tool baseline (NGMO)
Web search off; a research construct, not a training-data readout
600 responses, Sept 2
600Prompt Pairs
Each Run
Twice
CSPS Protocol
LSS Mode
Optional search (AGSO)
Web search enabled; invocation not logged, 83/600 cited
600 responses, Sept 2
ΔM₁ Output
Amplified · Stable · Suppressed
Original: 25 Stable · Forced: 13 / 3 / 9
Finding In the original optional-search condition, all 25 programs are Stable under the BH rule. Because the same 600 prompts ran in both modes, the two mention rates are paired: prompts where both modes agree carry no directional information, and only the discordant prompts do. The exact McNemar test on those pairs, corrected across 25 comparisons, found five programs at nominal p < 0.05, all in the suppression direction: McCombs (p 0.0033, q 0.0505), Ross (p 0.0040, q 0.0505), Tuck (p 0.0093, q 0.0664), Johnson (p 0.0106, q 0.0664), and Wharton (p 0.0201, q 0.1004). None survived correction. Stable means the difference did not reach the significance threshold in this sample. Whether the two configurations are equivalent is a separate question that this test leaves open. Nor does a low citation rate force ΔM₁ toward zero: the Persona stratum had zero citations in this condition, yet Stanford GSB was named in 0.800 of Persona optional-search responses versus 0.617 of no-tool responses, a difference of +0.183. Citation presence (13.8% overall) is a proxy for retrieval, not a logged invocation rate; where retrieval was not confirmed, the contrast is real but cannot be attributed to live search. The required-search condition, below, addresses that gap by logging invocation on every response.

Original optional-search ΔM₁ for all 25 programs, sorted. Every point sits inside ±0.05; the largest are Tuck and Ross at −0.045 and Stanford at +0.027. The ±0.05 lines mark the legacy Option D threshold for reference only. Statistical labels come from the paired McNemar/BH test, not from this gate, and no paired-interval materiality assessment is supplied in this release.

Optional search vs. no-tool baseline, September 2, 2026. Dotted lines mark the legacy ±0.05 Option D threshold, shown for reference.
Search-amplified direction Search-suppressed direction ±0.05 legacy Option D threshold (reference)
Program M₁ᴬᴳˢᴼ optional search M₁ᴺᴳᴹᴼ no-tool ΔM₁ |ΔM₁| BH class
Stanford GSB0.6270.600+0.0270.027Stable
Wharton0.6250.650−0.0250.025Stable
Harvard Business School0.6030.605−0.0020.002Stable
MIT Sloan0.5580.573−0.0150.015Stable
Columbia Business School0.5920.590+0.0020.002Stable
Kellogg0.5730.598−0.0250.025Stable
Haas0.4730.495−0.0220.022Stable
Booth0.4420.442+0.0000.000Stable
Tuck0.2800.325−0.0450.045Stable
Yale SOM0.2750.268+0.0070.007Stable
Stern0.2450.250−0.0050.005Stable
Ross0.1880.233−0.0450.045Stable
Fuqua0.1770.190−0.0130.013Stable
Darden0.1220.122+0.0000.000Stable
Anderson0.1170.130−0.0130.013Stable

Top 15 by original optional-search M₁ shown; all 25 are Stable in this condition. Full table in the Technical Appendix. Forced-search results follow in the required-search section.

What this means In the original comparison, optional search and the no-tool baseline produced similar mention rates for every program: the largest differences were under five percentage points and none reached BH significance. That is a description of two generated outputs on one date, not evidence that the configurations are interchangeable or that content aimed at one would carry to the other. An earlier version of this page reported a MIT Sloan Factual-stratum divergence (0.200 vs. 0.067) as a signal worth watching; that figure was an entity-detection artifact. The corrected MIT Sloan Factual rate is 0.050 in both modes (3 of 60), the same as the other top-six programs; the registry names MIT Sloan directly in three Factual prompts, so this is a named-lookup stratum, not evidence about spontaneous discovery or exceptional factual retrieval. The required-search comparison below is the more informative one: with invocation logged on every response, 13 programs were named significantly less often than in the reused baseline.
Act 2 How mentions vary by prompt type
Cross-stratum evenness · original optional-search condition

Cross-Stratum Consistency (CSC)

CSC = 1 minus CV(stratum M₁ᴬᴳˢᴼ scores). Range [0,1]. Higher = more even mention rate across the four designed strata (Generic, Constraints, Persona, Factual); lower = more uneven. Because the Factual stratum consists of named-school lookups with unequal coverage, CSC describes unevenness across the strata as designed, not established fragility or broad discovery consistency.

Finding McCombs has the highest CSC at 0.52 despite a low headline score: its low mention rates are the most even across the four strata. An earlier version of this page reported MIT Sloan as the CSC leader at 0.63 with a distinct Factual-stratum lift; that was an artifact of an entity-detection bug and does not hold under corrected v1.2.0 scoring, where MIT Sloan's CSC is 0.48 and its Factual stratum is unremarkable. Foster (CSC 0.00) appears in Generic (0.100) and Constraints (0.072) responses and in neither Persona nor Factual responses; Johnson (CSC 0.06) is concentrated in Generic (0.100), with a Factual rate of 0.050 from its three named prompts.
Program CSC Generic Constraints Persona Factual Profile
McCombs0.52 0.0880.0560.0170.050Consistent low
Kenan-Flagler0.43 0.0500.0110.0170.050Consistent low
Haas0.49 0.5170.5330.5080.050Factual drop-off
Booth0.48 0.5500.4440.4170.050Generic-led
MIT Sloan0.48 0.5790.6830.5830.050Constraints-led
Wharton0.48 0.6580.7170.7080.050Factual drop-off
Harvard Business School0.47 0.6000.7280.7000.050Factual drop-off
Columbia Business School0.47 0.6830.6890.5330.050Factual drop-off
Anderson0.45 0.1170.1830.0500.050Constraints-led
Kellogg0.46 0.5460.6670.7500.050Persona-led
Stanford GSB0.46 0.6000.7390.8000.050Persona-led
Marshall0.42 0.1000.0890.0080.050Generic-led
Fuqua0.41 0.1040.2280.3080.050Persona-led
Darden0.39 0.1540.0440.2080.050Generic/Persona
Tuck0.38 0.4330.1560.2750.050Generic-led
What this means A program with high M₁ but low CSC has visibility that varies sharply across the four designed strata. Stanford GSB's optional-search Persona rate of 0.800 (highest in the corpus) sits alongside a Factual rate of 0.050, the same Factual rate as every other top-six program. The Factual stratum consists of named-school lookup prompts (three per school for 19 tracked programs), so its rates reflect prompted lookups rather than spontaneous discovery. These stratum differences show that an overall rank can hide variation across query types; how often applicants use each type, and whether content aimed at a weaker stratum moves a program's rate there, are the next hypotheses to test. Under forced search, every top-six program's Constraints and Persona rates fell sharply while Factual stayed at 0.050 (see the forced-search strata in the appendix data).
Sentiment analysis

Valence: keyword-proxy framing (pooled no-tool and optional-search responses)

Valence scored via keyword-pattern matching in ±200-character context windows around the first matched occurrence for each program in each response. Positive = laudatory framing or list-position signal. Negative = critical or cautionary language.

0%
Negative mentions, corpus-wide
Zero negative framing detected across all 1,200 responses and 25 programs
40.9%
Positive mention ratio
59.1% neutral. Pooled across all 1,200 original responses (7,730 entity-response observations). "Ranked," "leading," and "prestigious" language dominate
0.644
Foster: highest valence ratio
47 positive of 73 total mentions. Small n amplifies ratio.
0.640
Harvard Business School valence
464 positive of 725 total, the largest absolute positive-mention count
Finding The keyword-based proxy found no negative-framing matches across the 1,200 September 2 responses. Zero matches mean the detector found none of the words on its negative list; semantic validation of framing, and generalization beyond the measured condition, are the next steps. Programs that appear were not contextualized negatively by this proxy; programs that do not appear are simply not named. Valence was not rescored for the required-search condition.
Positive vs. neutral framing for the 15 most-visible programs (by optional-search M₁), sorted by positive share; the pooled corpus is all 1,200 original responses. Foster, outside this top 15 by visibility, has the highest positive share in the cohort (0.64). Negative framing is 0% for every program, so that segment is omitted rather than drawn as an invisible sliver.
Positive framing Neutral framing
What this means Positive share is a ratio of each program's own mentions, so it is reported here as an observed framing proportion, not as a function of visibility: Foster, with one of the lowest mention rates, has the highest positive share in the cohort (0.64). Whether framing and visibility are related is an open question in this corpus. The zero-negative finding reflects the keyword proxy's behavior on this corpus, not a validated semantic characterization.
Act 3 What this means for strategy
Framework

AEO vs. SEO: why AI visibility is a distinct discipline

Search Engine Optimization (SEO) concerns how a page ranks in a results list. AI Engine Optimization (AEO) concerns how an entity is represented in generated answers. They measure different outcomes while sharing much of the same technical and content foundation; retrieval-based AEO still depends on search infrastructure.

The core distinction MBA-AEPI measures the AEO outcome only: whether a program is named in GPT-4o's answers under specified configurations. No SEO metrics are inputs to the index. The index measures outputs: a low mention rate locates where a program is absent from answers, and the cause, whether training coverage, web authority, or prompt fit, is the next thing to investigate. Content accessibility, clear entity descriptions, and source quality are the plausible levers for retrieval-based systems; testing them is the natural follow-on study.
Interpretation

Reading a program's profile: investigation paths, not diagnoses

A classification label describes a measured relationship between two configurations on the dates collected. Stable means the difference did not reach the BH significance threshold in this sample; equivalence is a separate question the test leaves open. The patterns below are places to look next. None is an intervention this study tested.

Observed patternPrograms (original optional search)What to investigate first
High M₁, Stable originally, suppressed under forced searchStanford, Wharton, HBS, Columbia, Kellogg, MIT SloanReplicate the drop with both conditions collected close together and the model snapshot pinned. A large gap between runs can reflect configuration, timing, sampling variation, or retrieval; it does not by itself establish dependence on one information source.
Mid M₁, low CSCYale SOM, Tuck, BoothReview the actual constraint and factual questions and answers. The program may not meet those constraints, its attributes may be unclear, or the detector may miss mentions. Content coverage is one hypothesis to test, not a conclusion from the gap alone.
Low M₁ in both configurationsMendoza, Olin, Carey, SmithFirst validate entity detection and prompt relevance, then check whether accurate, accessible information about the program exists. A low mention rate locates an absence; the cause is the next thing to investigate.
Significant but small increase under forced searchCarey (+0.030), Foster (+0.028), Smith (+0.012)Inspect the returned citations and the answers' factual claims. Each increase is BH-significant but under five percentage points and has not passed a separate materiality assessment.
Act 4 Required search
Required search · September 5, 2026 UTC

What changed when search was required

The required-search condition reran the same 600 prompt IDs with the search tool set to required, so that web search was invoked on every call, and logged invocation per response. It reused the original September 2 no-tool baseline; no new baseline was collected alongside it. The comparison therefore describes observed differences between two runs three days apart under an unpinned gpt-4o model alias. The recovered collection scripts also differ in one request setting: the original defaults to a 1,500-token output cap (overridable on the command line; the historical invocation is not archived), while the forced-search script sets no explicit cap. It does not fully isolate retrieval from timing, model-version, or request-setting differences.

600/600
Invocation flag set
Set when a response contains a web_search_call item or a citation annotation; not a retained raw tool-call trace. 17 flagged responses have no citations.
97.2%
Responses with citations
583 of 600; mean 6.79 citation annotations per response, not deduplicated unique sources, so not directly comparable to the original extractor's URL-deduplicated counts.
13 · 3 · 9
Suppressed · amplified · Stable
Paired exact McNemar + Benjamini-Hochberg, 25 programs, α = 0.05
0.447
Wharton: #1 under forced search
Stanford, #1 in optional search at 0.627, is #5 here at 0.383
Finding Requiring search produced lower mention rates for most programs relative to the reused no-tool baseline. Twenty of 25 programs had negative ΔM₁; 13 were significantly lower after BH correction: Harvard, Wharton, Columbia, Stanford GSB, Tuck, MIT Sloan, Kellogg, Haas, Yale SOM, Fuqua, Ross, Booth, and Anderson. The largest difference was Columbia at −0.230 (0.590 → 0.360). Three programs were significantly higher: Carey (+0.030), Foster (+0.028), and Smith (+0.012), each under five percentage points. Nine were Stable. Of the five nominal suppression candidates from the original condition, three (Wharton, Tuck, Ross) were significant here; McCombs and Johnson were Stable. These are statistical classifications; effect sizes are reported alongside them. The separate paired-interval materiality assessment described in Amendment v1.9.1 has not been performed.

Citation presence by prompt type, original optional search vs. forced search. The Persona stratum, which had no citations in the original condition, produced cited answers in 115 of 120 forced-search responses. That confirms these prompts can yield retrieval; it does not retrospectively establish whether search ran in the original uncited answers.

StratumOptional search, with citationsForced search, with citationsForced search, invocation logged
Generic14 / 240 (5.8%)235 / 240 (97.9%)240 / 240
Constraints19 / 180 (10.6%)174 / 180 (96.7%)180 / 180
Persona0 / 120 (0.0%)115 / 120 (95.8%)120 / 120
Factual50 / 60 (83.3%)59 / 60 (98.3%)60 / 60
All strata83 / 600 (13.8%)583 / 600 (97.2%)600 / 600

Forced-search ΔM₁ for all 25 programs, sorted. Compare the scale with the original chart above: the largest original difference was −0.045, while here 12 programs fall past −0.05 (the legacy Option D threshold, shown for reference) and a thirteenth, Anderson at −0.045, is BH-significant without crossing it. Labels come from the paired McNemar/BH test.

Forced search (Sept 5 UTC) vs. the reused Sept 2 no-tool baseline. Dotted lines mark the legacy ±0.05 threshold, for reference only.
Search-Amplified (BH) Search-Suppressed (BH) Stable (BH) ±0.05 legacy threshold (reference)

The forced-search ranking. Ranked on M₁ᴬᴳˢᴼ under required search. "Rank, optional" is the same program's rank in the original optional-search condition; "Change" is optional rank minus forced rank. Columbia and Booth tie at 0.360 and share rank 7.

Rk Program M₁ᴺᴳᴹᴼ no-tool M₁ forced search ΔM₁ Rank, optional Change BH class
1Wharton0.6500.447−0.2032+1Search-Suppressed
2Kellogg0.5980.408−0.1905+3Search-Suppressed
3Harvard Business School0.6050.387−0.21830Search-Suppressed
4MIT Sloan0.5730.385−0.1886+2Search-Suppressed
5Stanford GSB0.6000.383−0.2171-4Search-Suppressed
6Haas0.4950.382−0.1137+1Search-Suppressed
7Booth0.4420.360−0.0828+1Search-Suppressed
7Columbia Business School0.5900.360−0.2304-3Search-Suppressed
9Stern0.2500.235−0.01511+2Stable
10Yale SOM0.2680.188−0.080100Search-Suppressed

Top 10 by forced-search M₁ shown. All 25, with confidence intervals and the legacy Option D label alongside the primary BH label, are in the Technical Appendix.

What this means Several programs with high baseline mention rates were named less often when search was required, and the leaderboard reordered: Stanford GSB moved from first to fifth, Columbia from fourth to a tie for seventh, Kellogg from fifth to second. Because the baseline was collected three days earlier under a model alias that may have changed, these are measured differences between two collections; a concurrent baseline in the next collection will reduce the timing confound; isolating retrieval itself also requires matched request settings, model-version controls, and retained invocation evidence. What the required-search condition does resolve is the invocation question: search demonstrably ran on every response, and Persona prompts demonstrably can produce cited answers. The original condition's invocation status remains unknown. Future dual-mode collections should pair a no-tool baseline with required search, collected close together, with invocation logging and a pinned model snapshot.
Post-hoc robustness check · 540 discovery-style prompts

Did the findings depend on the named-school factual prompts?

Because all 60 Factual prompts name schools, we excluded them symmetrically from both arms of each comparison and repeated the analysis on the remaining 540 prompt IDs (240 Generic, 180 Constraints, 120 Persona; effective weights 44.4% / 33.3% / 22.2%, no reweighting). Regex detection is unchanged; Wilson intervals use n = 540; exact McNemar tests and BH correction are rerun within each comparison. This is a sensitivity analysis, not a new pre-specified index, and the two analyses share responses, so it is not an independent replication.

Finding Excluding the 60 named-school factual prompts preserved the principal results. The original optional-search ranking was unchanged for all 25 programs (Spearman 1.0), and all 25 remained Stable. Under forced search the top six programs and every program-level BH classification were unchanged: 13 lower, 3 higher, 9 Stable (rank agreement with the 600-prompt analysis: Spearman 0.9985). Three lower-table forced-search ranks moved (Tepper 20 → 18; McCombs 19 → 20; Kenan-Flagler 22 → 23). Mention rates rise mechanically because the low-mention Factual stratum leaves the denominator; these are not improvements over time. Citation presence falls to 33/540 (6.1%) for optional search and 524/540 (97.0%) for forced search, following directly from removing the highly cited Factual stratum. This check shows the headline pattern did not hinge on the named-school prompts in these saved responses. It does not make the 600-prompt design entity-free, validate detection against human labels, or remove the timing and request-setting differences.
RkProgram (forced-search top six)M₁, 600 promptsM₁, 540 promptsΔ vs. no-tool, 600Δ vs. no-tool, 540BH class (both)
1Wharton0.4470.491-20.33 pp-22.59 ppSearch-Suppressed
2Kellogg0.4080.448-19.00 pp-21.11 ppSearch-Suppressed
3Harvard Business School0.3870.424-21.83 pp-24.26 ppSearch-Suppressed
4MIT Sloan0.3850.422-18.83 pp-20.93 ppSearch-Suppressed
5Stanford GSB0.3830.420-21.67 pp-24.07 ppSearch-Suppressed
6Haas0.3820.419-11.33 pp-12.59 ppSearch-Suppressed

Largest negative difference on 540 prompts remains Columbia (-25.56 pp); the three amplified programs remain Carey (+3.33 pp), Foster (+2.96 pp), and Smith (+1.30 pp). Only Foster and Smith have changed BH q values (0.0305 and 0.0260), both still below 0.05; all optional-search p/q values are identical because the excluded Factual observations contributed no discordant pairs. Source: sensitivity/comparison_600_vs_540.csv and sensitivity/summary.json.

Release scope This page covers ChatGPT (GPT-4o API) only, under three tested configurations. Results for other engines and any cross-engine comparison are outside this release and are not reported here. Next steps for this instrument: replicate the forced-search comparison with closely timed conditions, explicit invocation logging, a pinned model snapshot, and human-validated entity detection. Disclosure: Arcalea offers AEO consulting services. This study measures AI visibility: how often GPT-4o, accessed through the OpenAI API, names each MBA program under three specified search configurations, reported alongside the U.S. News quality ranking.
Infrastructure How this is scored
Protocol

Methodology: v1.9.0-FROZEN with Amendment v1.9.1

The no-tool and optional-search conditions each collected one response per prompt for 600 prompts in one session on September 2, 2026 (CSPS protocol). The required-search condition is a documented exception: its 600 responses are timestamped September 5 UTC and are compared against the reused September 2 baseline, so prompt pairing controls the question asked but not the collection date. Statistical labels follow Amendment v1.9.1: paired exact McNemar with Benjamini-Hochberg correction across 25 programs; the legacy Option D rule is retained separately.

Prompt registry: a correction The recovered 600-prompt registry is not entirely entity-free, as earlier descriptions stated. The 540 Generic, Constraints, and Persona prompts passed an automated scan for target-school names. All 60 Factual prompts (MBA-PROD-0541 to 0600) explicitly name schools: three question variants for each of 20 schools, 19 of them in the scored cohort and one (Emory Goizueta) outside it. Six tracked programs receive no direct factual question. All 600 stored SHA-256 hashes match the prompt text, so this is a design and documentation issue, not a data-integrity failure. Every aggregate score on this page is therefore a mixed discovery/lookup index and should not be read as wholly spontaneous discovery. The numerical results are unchanged; a separately labeled 540-prompt robustness check is reported in the sensitivity analysis.
600
Prompt pairs
240 Generic · 180 Constraints · 120 Persona · 60 Factual
1,800
Captures, two collections
600 no-tool + 600 optional search (Sept 2) + 600 forced search (Sept 5 UTC)
25
Target programs
All 25 detected ≥1 response · 0 zero-mention programs
Core scoring formulas

M₁ᴬᴳˢᴼ = mentions_search / 600: search-enabled mention rate, labeled by condition

M₁ᴺᴳᴹᴼ = mentions_no-tool / 600: no-tool baseline mention rate

ΔM₁ = M₁ᴬᴳˢᴼ − M₁ᴺᴳᴹᴼ: configuration contrast

CSC = max(0, 1 − σ/μ): consistency across 4 strata (σ=SD, μ=mean)

Wilson CI: (p̂ + z²/2n ± z·√(p̂(1−p̂)/n + z²/4n²)) / (1 + z²/n): α=0.05, z=1.96

Primary label: exact McNemar on discordant pairs (b = no-tool only, c = search only) → BH q across 25; Amplified if q ≤ 0.05 and ΔM₁ > 0, Suppressed if q ≤ 0.05 and ΔM₁ < 0, else Stable

Legacy Option D (diagnostic only): |ΔM₁| ≥ 0.05 AND non-overlapping Wilson CIs

Technical appendix

Data governance & complete entity tables

Table A: original optional-search condition (September 2, 2026). Ranked on M₁ᴬᴳˢᴼ with optional search. Source: mba_aepi_index_scores_wave1_chatgpt_v2.csv. Table B, the forced-search condition, follows.

Rk Program M₁ᴬᴳˢᴼ 95% CI M₁ᴺᴳᴹᴼ ΔM₁ CSC Valence Mentions
1Stanford GSB0.627[0.587, 0.664]0.600+0.0270.460.62736
2Wharton0.625[0.586, 0.663]0.650−0.0250.480.59765
3Harvard Business School0.603[0.564, 0.642]0.605−0.0020.470.64725
4Columbia Business School0.592[0.552, 0.630]0.590+0.0020.470.37709
5Kellogg0.573[0.533, 0.612]0.598−0.0250.460.27703
6MIT Sloan0.558[0.518, 0.598]0.573−0.0150.480.33679
7Haas0.473[0.434, 0.513]0.495−0.0220.490.33581
8Booth0.442[0.402, 0.482]0.442+0.0000.480.32530
9Tuck0.280[0.246, 0.317]0.325−0.0450.380.32363
10Yale SOM0.275[0.241, 0.312]0.268+0.0070.220.46326
11Stern0.245[0.212, 0.281]0.250−0.0050.430.36297
12Ross0.188[0.159, 0.222]0.233−0.0450.320.19253
13Fuqua0.177[0.148, 0.209]0.190−0.0130.410.25220
14Darden0.122[0.098, 0.150]0.122+0.0000.390.32146
15Anderson0.117[0.093, 0.145]0.130−0.0130.450.24148
16Tepper0.087[0.067, 0.112]0.097−0.0100.120.32110
17Marshall0.073[0.055, 0.097]0.078−0.0050.420.3391
18Foster0.062[0.045, 0.084]0.060+0.0020.000.6473
19McCombs0.060[0.044, 0.082]0.088−0.0280.520.3789
20Johnson (Cornell)0.048[0.034, 0.069]0.070−0.0220.060.3971
21Kenan-Flagler0.032[0.020, 0.049]0.028+0.0030.430.2836
22Mendoza0.020[0.011, 0.035]0.018+0.0020.000.2623
22Olin (WashU)0.020[0.011, 0.035]0.020+0.0000.000.0824
24Carey (Johns Hopkins)0.015[0.008, 0.028]0.013+0.0020.000.5317
25Smith (Maryland)0.013[0.007, 0.026]0.012+0.0020.000.3315

Rank 22 repeats and rank 23 is skipped, not a typo: Mendoza and Olin are exactly tied at M₁ᴬᴳˢᴼ = 0.020 (12 of 600 responses each) and share rank 22, so the next distinct value (Carey, 0.015) takes rank 24. Standard competition-ranking convention.

Source, Table A: MBA-AEPI · ChatGPT GPT-4o (alias, snapshot not recorded) · September 2, 2026 · 1,200 captures · 100% completion · response length 207 to 5,964 characters (median 1,588.5). All 25 programs are Stable under the BH rule in this condition.

Table B: forced-search condition

Ranked on M₁ᴬᴳˢᴼ with the search tool set to required, collected September 5, 2026 UTC, against the reused September 2 no-tool baseline. "BH class" is the primary label (paired exact McNemar, Benjamini-Hochberg across 25, α 0.05). "Legacy Option D" is the older CI-overlap/five-point diagnostic, shown for continuity only. Source: mba_aepi_index_scores_wave1b_chatgpt_v3.csv.

Rk Program M₁ forced 95% CI M₁ᴺᴳᴹᴼ ΔM₁ Rk, optional BH class Legacy Option D
1Wharton0.447[0.407, 0.487]0.650−0.2032Search-SuppressedSearch-Suppressed
2Kellogg0.408[0.370, 0.448]0.598−0.1905Search-SuppressedSearch-Suppressed
3Harvard Business School0.387[0.349, 0.426]0.605−0.2183Search-SuppressedSearch-Suppressed
4MIT Sloan0.385[0.347, 0.425]0.573−0.1886Search-SuppressedSearch-Suppressed
5Stanford GSB0.383[0.345, 0.423]0.600−0.2171Search-SuppressedSearch-Suppressed
6Haas0.382[0.344, 0.421]0.495−0.1137Search-SuppressedSearch-Suppressed
7Booth0.360[0.323, 0.399]0.442−0.0828Search-SuppressedSearch-Suppressed
7Columbia Business School0.360[0.323, 0.399]0.590−0.2304Search-SuppressedSearch-Suppressed
9Stern0.235[0.203, 0.271]0.250−0.01511StableStable
10Yale SOM0.188[0.159, 0.222]0.268−0.08010Search-SuppressedSearch-Suppressed
11Ross0.148[0.122, 0.179]0.233−0.08512Search-SuppressedSearch-Suppressed
12Tuck0.133[0.108, 0.163]0.325−0.1929Search-SuppressedSearch-Suppressed
13Darden0.118[0.095, 0.147]0.122−0.00314StableStable
14Fuqua0.115[0.092, 0.143]0.190−0.07513Search-SuppressedSearch-Suppressed
15Johnson0.093[0.073, 0.119]0.070+0.02320StableStable
16Foster0.088[0.068, 0.114]0.060+0.02818Search-AmplifiedStable
17Anderson0.085[0.065, 0.110]0.130−0.04515Search-SuppressedStable
18Marshall0.082[0.062, 0.106]0.078+0.00317StableStable
19McCombs0.080[0.061, 0.104]0.088−0.00819StableStable
20Tepper0.077[0.058, 0.101]0.097−0.02016StableStable
21Carey0.043[0.030, 0.063]0.013+0.03024Search-AmplifiedStable
22Kenan-Flagler0.023[0.014, 0.039]0.028−0.00521StableStable
22Smith0.023[0.014, 0.039]0.012+0.01225Search-AmplifiedStable
24Olin0.017[0.009, 0.030]0.020−0.00322StableStable
25Mendoza0.005[0.002, 0.015]0.018−0.01322StableStable

Ranks 7 and 22 repeat: Columbia and Booth tie at 0.360; Smith and Kenan-Flagler tie at 0.023. Competition-ranking convention. Legacy Option D yields 12 suppressed / 13 Stable; it is not the headline classification.

Source, Table B: MBA-AEPI · ChatGPT GPT-4o · September 5, 2026, 01:14 to 02:16 UTC · 600 captures · invocation flag 600/600 · citations in 583/600 (97.2%), mean 6.79 annotations per response. Provenance: both raw ChatGPT collections (1,200 original responses timestamped Sept 2, 16:59 to 18:49 UTC; 600 forced-search responses), the 600-prompt registry with SHA-256 hashes, and both collection scripts are retained in the research repository and fingerprinted in chatgpt_source_manifest.json; they are not included in the downloads on this page, which are limited to the methodology and companion documents. The read-only verifier reconstructs every mention count, both paired-count matrices, all p/q values, and all 600 registry hashes from the raw text. Open items it leaves for future work are detector accuracy against human labels, the exact historical command-line settings or SDK versions, and capture-time hash validation, which the recovered scripts did not perform.

Common questions about MBA-AEPI

What is MBA-AEPI?

MBA-AEPI (MBA AI Engine Perception Index) measures how often GPT-4o, accessed through the OpenAI API, names each of 25 MBA programs in answers to a fixed 600-prompt registry under specified configurations. This ChatGPT release covers three configurations: a no-tool baseline and an optional-search condition (September 2, 2026) and a required-search condition (September 5 UTC), 1,800 responses in total. It measures AI visibility, how often each program is named in generated answers under those configurations, reported alongside established quality rankings.

Are the prompts entity-free?

No, not all of them. The 540 Generic, Constraints, and Persona prompts passed an automated scan for target-school names, but all 60 Factual prompts explicitly name schools (three variants each for 20 schools, 19 of them tracked). Earlier descriptions calling every prompt entity-free were incorrect. The aggregate scores are therefore a mixed discovery/lookup index. Scores were not changed; a separate 540-prompt robustness check is reported on the page.

Did the findings depend on the named-school factual prompts?

In these saved responses, no. Excluding all 60 Factual prompts left the original optional-search ranking unchanged for all 25 programs, kept all 25 Stable, and left the forced-search top six and all 13 / 3 / 9 BH classifications unchanged, with three lower-table rank moves (Tepper, McCombs, Kenan-Flagler). Rates rise mechanically because the denominator shrinks. This is a post-hoc robustness check, not an independent replication.

Which ranking is “the” ChatGPT ranking?

There are two, and they differ. In the original optional-search condition Stanford GSB led at 0.627, followed by Wharton 0.625, Harvard 0.603, Columbia 0.592, Kellogg 0.573, and MIT Sloan 0.558. In the forced-search condition Wharton led at 0.447, followed by Kellogg 0.408, Harvard 0.387, MIT Sloan 0.385, and Stanford 0.383. Always label the condition; neither is simply the ChatGPT rank.

What does a Stable classification mean?

Stable means the program's mention-rate difference between two configurations did not reach significance under a paired exact McNemar test with Benjamini-Hochberg correction across 25 programs at alpha 0.05. Whether the configurations are equivalent is a separate question that this test leaves open. In the original optional-search condition all 25 programs were Stable; under forced search, 13 were Search-Suppressed, 3 Search-Amplified, and 9 Stable. Statistical significance is separate from practical materiality, which this release does not assess.

Did ChatGPT actually search the web?

For the original optional-search condition, invocation was not logged; citations appeared in 83 of 600 responses (13.8%), and citation presence is a proxy, not a measured invocation rate. The required-search condition set an invocation flag on all 600 responses; that flag is set by either a web_search_call item or a citation annotation and is not a retained raw tool-call trace. 583 responses (97.2%) contained citations. Zero citations do not imply a zero difference in mention rates: Stanford's Persona rate was 0.800 with optional search versus 0.617 without, despite no citations.

Can the results be reproduced?

Both raw ChatGPT collections, the 600-prompt registry with hashes, both collection scripts, the scoring code, and all outputs are retained in the research repository; they are not included in the downloads on this page, which are limited to the methodology and companion documents. A read-only verifier reconstructs every mention count, both paired comparisons, all p/q values, and all registry hashes from the raw text without API calls. Open items it leaves for future work are detector accuracy against human labels, the exact historical command-line settings or SDK versions, and future model outputs.

Why does this release cover ChatGPT only?

The release scope is ChatGPT under no-tool, optional-search, and required-search conditions. Results for other engines and any cross-engine comparison are outside this release and are not reported here. The next step for this instrument is to replicate the forced-search comparison with closely timed conditions, explicit invocation logging, a pinned model snapshot, and human-validated entity detection.

How is AI visibility different from a traditional MBA ranking?

A traditional ranking (US News, Financial Times) scores programs on inputs like selectivity, salary outcomes, and survey results. MBA-AEPI scores something else: how often GPT-4o names a program in answers to a fixed prompt set under a specified configuration. In the original optional-search condition that ranking tracked the recorded US News rank closely (Spearman rho 0.973), while individual programs still sat above or below the fitted line. The study does not inspect training data or identify which sources shaped an answer, so it does not explain why a program appears more or less often.

What does CSC mean?

CSC (Cross-Stratum Consistency) measures how evenly a program's mention rate holds across the four query strata: Generic, Constraints, Persona, and Factual. A high CSC means the program is named consistently across query types; a low CSC means its mention rates are uneven across the four designed strata. CSC = 1 minus the ratio of standard deviation to mean across the four stratum M1 values, clipped to [0, 1]. CSC values on this page are from the original optional-search condition; note that the Factual stratum is a named-school lookup stratum.