Built on a library of 105 primary documents spanning the NBER, IMF, OECD, BIS, Federal Reserve System, World Bank and the major AI labs and forecasters — ENSI Foresight Division.
The forecast that hides the thing it should be measuring
Every institutional forecast of AI’s economic effect — Goldman Sachs’ 7%, PwC’s $15.7 trillion, the OECD’s 0.25–0.6 percentage points, the IMF’s country-by-country growth differentials — answers a narrower question than the one policymakers think they are asking. Each is, in its own methodology, a model of potential output: what the economy could produce if the productivity gain shows up and is spent. None of the headline numbers that dominate boardroom slides and finance-ministry briefings is a model of realized, demand-constrained GDP — the number that determines whether tax receipts rise, whether unemployment claims fall, and whether a government gets re-elected. That gap between potential and realized output is not a technicality. It is the single most consequential and most underpriced risk in the entire AI-and-growth literature, and this report exists to make it visible.
The reframe this report argues for is simple to state and easy to miss: the modal, mainstream forecast — modest, positive, uneven aggregate growth — is entirely compatible with regions, sectors and income deciles inside that aggregate experiencing outright contraction, for the same reason and at the same time. A national GDP print of +0.4% is arithmetically consistent with call-centre towns, back-office cities and mid-skill service corridors losing income for a decade, exactly as Pittsburgh, Greensboro and the furniture counties of North Carolina did during the “China Shock” — not because the aggregate model was wrong, but because aggregate models are built to net out precisely the geography and distribution that determines whether a shock feels like growth or like shrinkage to the people living through it. Aggregation is not measurement error. It is a modelling choice, and every one of the headline AI-growth forecasts in this library makes it.
This matters because the entities that most need an accurate scenario map — national treasuries, central banks, regional development agencies, sovereign wealth funds, multilateral lenders — are not making a bet on a single global GDP number. They are making a portfolio of decisions about fiscal buffers, retraining budgets, monetary stance, regional transfers and industrial policy, each of which depends on knowing not just whether AI grows the pie but whose slice moves and how fast the redistribution happens relative to the political cycle that has to absorb it. A state that plans against Goldman Sachs’ 7% uplift and gets the OECD’s 0.4 percentage points instead has made a forecasting error. A state that plans against any aggregate figure at all, while a fifth of its regions or a third of its labour force experiences the Korinek-Stiglitz demand-shrinkage mechanism described in Scenario 4 below, has made a category error — and it is the more dangerous of the two, because the national dashboard will keep reading “growth” the entire time.
The five-scenario framework that follows is deliberately not a spectrum from “big growth” to “big shrinkage” with the truth somewhere in the middle. It is four distinct causal stories — a productivity boom, a modest base case, a stagnation trap, and a demand-driven contraction — each anchored to a different published model with a different methodology and a different set of assumptions about diffusion speed, gain-sharing and constraint-bindingness, plus a fifth argument that cuts across all four: that the geographic and sectoral variance inside whichever aggregate number wins is where the real policy risk lives. Readers should leave this report able to name, for their own jurisdiction, which leading indicators would move them from one scenario to another — and understanding that the question “will AI grow or shrink the economy” is underspecified until you also ask “for whom, and measured how.”
One discipline runs through every section below: history’s base rate for how long general-purpose technologies take to resolve into measured productivity. Paul David’s canonical study, “The Dynamo and the Computer,” published in the American Economic Review, found that electrification took roughly forty years to show up in US productivity statistics after the technology existed — because factories had to be rebuilt around unit-drive motors rather than retrofitted around a single central power source, and that reorganisation, not the invention, was the rate-limiting step. Nicholas Crafts’ growth-accounting study for the LSE, “Steam as a General Purpose Technology,” pushes the base rate further: steam power took close to a century to lift UK aggregate productivity growth measurably. Every probability estimate in this report should be read against that discipline. A scenario “resolving” by 2030 or even 2035 would be fast by the standard of the only two general-purpose technologies we have full-century data on.
The five futures, at a glance
The Productivity Boom (~15% likelihood). Fast diffusion, broadly shared gains, no binding physical constraint — the world Goldman Sachs and PwC model. Attractive, well-funded, and the least likely of the four core scenarios because it requires three separate things the historical and structural evidence argues against simultaneously.
The Base Case — Modest, Uneven Growth (~45–50% likelihood). The mainstream institutional consensus: Acemoglu’s task-based ceiling, the OECD’s general-equilibrium range, the IMF and BIS’s cross-country findings. Growth that is real, arrives slowly, and is unevenly distributed by construction — the most probable single outcome and the one this report spends the most time defending.
Stagnation / The Productivity Paradox Redux (~15–20% likelihood). Diffusion stalls the way it stalled for computers in the 1980s and 1990s; the AI capex boom outruns realized value the way the BIS’s contest-theory model predicts; Baumol’s cost disease caps what automation alone can deliver.
Demand-Driven Shrinkage (~10–15% likelihood). The scenario readers most need explained because it is the least intuitive: potential output rises even as realized GDP falls, because the same productivity shock that looks expansionary in a supply-side model becomes contractionary once you close the model with a demand side that depends on wage income households don’t have.
The Bifurcation (the report’s central reframe, not a fifth probability bucket). A “modest growth” aggregate (Scenario 2) can mechanically conceal regional and sectoral versions of Scenario 4 playing out underneath it — because the published aggregate models are not built to catch that. This is the sharpest, most underpriced risk in the whole literature, and it is why “which scenario wins nationally” is the wrong question for any policymaker to be asking alone.
Scenario 1 — The Productivity Boom (~15% likelihood)
The bull case is not a fringe position; it is argued by two of the most-cited institutions in applied macro forecasting, and its arithmetic is worth taking seriously before it is discounted. Goldman Sachs’ Joseph Briggs and Devesh Kodnani, in “The Potentially Large Effects of Artificial Intelligence on Economic Growth,” estimate that roughly two-thirds of current US and European jobs are exposed to some degree of AI automation, that generative AI could substitute for up to one-quarter of current work tasks, and that extrapolated globally this is equivalent to exposing the labour input of some 300 million full-time jobs to automation. Run through their model, that yields just under 1.5 percentage points of additional annual US labour-productivity growth over a ten-year period following widespread adoption, and an eventual 7% increase in global GDP — about $7 trillion in their 2023 pricing. PwC’s “Sizing the Prize” is, if anything, the more dramatic of the two headline forecasts in this library: $15.7 trillion of additional global GDP by 2030, equivalent to global output being up to 14% higher than it would otherwise be, split by PwC’s own accounting into productivity effects (businesses automating processes and augmenting labour, which the firm says account for over 55% of total GDP gains between 2017 and 2030) and a second, growing consumption-side effect as AI-enhanced products and personalisation drive additional demand — a channel PwC says will account for 58% of the GDP gain realized in 2030 specifically, i.e. the mix shifts from productivity-led to demand-led as the diffusion matures. PwC further disaggregates by geography and sector: China stands to see the largest proportional boost (up to 26% of GDP), North America the second largest (14%), and retail, financial services and healthcare the sectors with the greatest combined productivity-and-product-enhancement potential.
Why does ENSI Foresight Division treat this as the low-probability tail rather than the central case, despite the credibility of the institutions behind it? Three reasons, each traceable to a different angle of this library, and each a condition the Boom scenario requires to hold simultaneously — which is precisely what makes their joint probability low even when each is individually plausible.
It requires diffusion speed that the historical record argues against. Paul David’s electricity study and Nicholas Crafts’ steam study — the two general-purpose technologies with a full-century data record — took 40 and roughly 100 years respectively to show up in aggregate productivity, because the complementary reorganisation of firms, not the invention itself, was the binding constraint. The Federal Reserve Board’s 2025 working paper “Generative AI at the Crossroads: Light Bulb, Dynamo, or Microscope?” tests generative AI explicitly against these historical diffusion analogues and finds the honest answer is not yet knowable — the paper’s title is itself an admission that AI could be any of the three, with very different growth implications.
It requires broad gain-sharing that the market-structure evidence directly contradicts. Jan De Loecker and Jan Eeckhout’s NBER working paper “The Rise of Market Power and the Macroeconomic Implications” documents that average US markups rose from roughly 18% above marginal cost in 1980 to 67% today, concentrated in an increase in high-markup firms rather than a broad-based shift — and they link this rise causally to falling labour share, falling low-skill wages, and a slowdown in aggregate output. Autor, Dorn, Katz, Patterson and Van Reenen’s companion NBER paper, “The Fall of the Labor Share and the Rise of Superstar Firms,” finds the same winner-take-most dynamic in industry concentration data. If AI’s productivity gains flow disproportionately to the handful of firms that already control the compute, the cloud infrastructure and the frontier models — exactly the concentration the UK Competition and Markets Authority’s “AI Foundation Models” technical report and the FTC’s 6(b) study of cloud-AI partnerships were opened to investigate — then the Boom scenario’s assumption of broadly shared productivity gains breaks down at the first link in the chain.
It requires no binding compute or energy constraint, which Angle 04’s evidence says is not a safe assumption. The IEA’s “Energy and AI” special report models data-centre electricity demand as a potential bottleneck on how fast AI compute can scale at all — grid capacity, transformer lead times and permitting are not software problems that can be diffused at ChatGPT’s adoption speed. The BIS’s “The AI Investment Race” (Phurichai Rungcharoenkitkul, July 2026) goes further and models the entire AI buildout as a winner-take-most contest in which competing firms rationally over-commit resources: calibrated to balance-sheet and deal data, the paper finds over-investment running at roughly 1.5 times the efficient level, rising to around 3 times where demand is less elastic — financed substantially through debt and circular equity ties between hyperscalers, model labs and chip suppliers, with a network-cascade risk if any one node disappoints on revenue.
None of this means the Boom scenario is impossible — Goldman and PwC’s methodologies are serious and their authors are not naïve about diffusion lags. It means the scenario requires fast diffusion, broad sharing and unconstrained capital deepening to all hold at once, against a historical base rate, a market-structure trend and a physical-infrastructure constraint that each independently argue the opposite. Compounding three low-conditional-probability requirements is why ENSI Foresight Division prices this scenario at roughly 15% — high enough to take seriously in capital allocation and infrastructure planning, low enough that a state should not build its ten-year fiscal plan on it arriving.
Scenario 2 — The Base Case: Modest, Uneven Growth (~45–50% likelihood)
This is the mainstream institutional consensus, and it deserves to be argued for on its own terms rather than treated as the residual “everyone else’s number.” Its intellectual anchor is Daron Acemoglu’s “The Simple Macroeconomics of AI,” prepared for Economic Policy and circulated as NBER Working Paper 32487. Acemoglu builds a task-based model in which AI’s macroeconomic effect is bounded by a version of Hulten’s theorem: aggregate productivity gains are given by the fraction of tasks AI actually touches multiplied by the average task-level cost saving. Using the best available exposure and productivity-improvement estimates, he finds the resulting TFP effect is “nontrivial but modest — no more than a 0.66% increase in total factor productivity over 10 years.” He then argues this may still be an overestimate, because current evidence is drawn disproportionately from easy-to-learn tasks, while much of AI’s future effect will have to come from hard-to-learn tasks with context-dependent judgment and no objective performance metric to train against — on that adjustment, his predicted 10-year TFP gain falls to under 0.53%. This is not a paper written to be contrarian for its own sake: Acemoglu explicitly engages Goldman’s 7% and McKinsey’s $17.1–25.6 trillion range in his introduction and argues the gap is explained by his more conservative estimate of which tasks are genuinely automatable versus merely exposed.
The OECD’s “Miracle or Myth? Assessing the Macroeconomic Productivity Gains from Artificial Intelligence” (OECD Artificial Intelligence Papers No. 29, November 2024) arrives independently at a strikingly similar order of magnitude through an entirely different method — a novel micro-to-macro multi-sector general-equilibrium model with input-output linkages, rather than Acemoglu’s task-exposure algebra. Its headline finding: annual aggregate TFP growth attributable to AI of 0.25–0.6 percentage points, equivalent to 0.4–0.9 percentage points of labour productivity growth over a ten-year horizon. That two independent methodologies — one a stylised task-based bound, one a full general-equilibrium simulation with sectoral input-output linkages — converge on the same order of magnitude, an order of magnitude roughly a fifth to a tenth the size of Goldman’s or PwC’s headline figures, is the single strongest piece of evidence in this library for treating the Base Case as the modal outcome rather than a competitor to the Boom scenario.
Ten to fifteen supporting reasons this is where ENSI Foresight Division places the largest probability mass:
It is where two independent, methodologically distinct institutional estimates converge, as above — Acemoglu’s task-exposure bound and the OECD’s general-equilibrium simulation, arrived at without coordination, land in the same 0.3–0.9 percentage-point-of-productivity-growth range.
It matches the observed pattern of real deployment, not just theory. The BIS’s “Artificial Intelligence and Growth in Advanced and Emerging Economies: Short-Run Impact,” a 56-economy, 16-industry cross-country study, finds a real but modest growth effect concentrated in advanced economies, consistent with a diffusion process still in its early innings rather than a step-change already realized.
It is consistent with firm-level field evidence showing real but bounded gains. Erik Brynjolfsson, Danielle Li and Lindsey Raymond’s “Generative AI at Work” studies 5,179 customer-support agents given access to a generative AI assistant and finds a 14% average productivity gain — economically significant, globally scalable in principle, but a 14% task-level gain is a different order of magnitude from a Goldman-style productivity boom, and the paper finds the gain is concentrated among novice and low-skilled workers (34% improvement) with minimal effect on already-experienced staff — a distributional pattern, not a uniform lift.
Software-development field experiments tell the same story of real, bounded, unevenly distributed gains. Microsoft Research’s three-firm RCT across 4,867 developers finds a 26% increase in task completion; the earlier GitHub Copilot field experiment finds developers completed a standardised task 55.8% faster. These are large individual-task effects that nonetheless net out, at the Acemoglu-style aggregate level, to a modest TFP contribution once weighted by the share of the economy actually composed of tasks this exposed.
The historical diffusion base rate argues for “modest and slow” over “large and fast.” The Productivity J-Curve literature (NBER Working Paper 25148) formalises why GPT adoption should be expected to show up first as a dip in measured productivity — as firms spend on complementary intangible investment that national accounts do not capture as capital — before any acceleration appears, which is exactly consistent with a base case that looks unremarkable for years before it looks real.
It is a growth effect real enough to matter for a finance ministry, wrong enough to disappoint an equity analyst pricing a Goldman-sized re-rating — which is itself diagnostic. Sustained modest TFP acceleration compounds: even Acemoglu’s more conservative 0.53% figure, sustained and extended past the initial ten-year window as diffusion continues per the historical base rate above, is not nothing over a multi-decade horizon.
It is unevenly distributed by construction, not by exception — the IMF’s “The Global Impact of AI: Mind the Gap” (WP/25/76) finds the estimated growth impact in advanced economies could be more than double that in low-income economies once sectoral exposure, technological preparedness and data/technology access are fed into a multi-region dynamic general-equilibrium model, with AI-driven productivity gains concentrated in the non-tradable sector large enough to disrupt the traditional exchange-rate-adjustment mechanism (an inverse Balassa-Samuelson effect, in the paper’s own framing).
It coexists with — rather than reverses — the concentration trend already documented in Angle 06. Nothing about a modest aggregate TFP gain requires that gain to be evenly shared; De Loecker and Eeckhout’s markup evidence and the superstar-firm literature describe a economy-wide trend already three decades in train, and the Base Case scenario simply assumes AI does not interrupt it, which is the more conservative and more defensible assumption than assuming AI reverses it.
It is what “Miracle or Myth?” itself frames as the honest middle — the OECD paper’s own Figure 1, comparing predicted macro-level productivity gains across the published studies, is explicitly built to show how much the estimates vary depending on methodology, and situates its own general-equilibrium estimate as the more disciplined, assumption-transparent number against which the use-case-based (McKinsey) and task-exposure (Goldman) estimates should be read as upper bounds rather than central forecasts.
It requires the fewest simultaneous strong assumptions. Unlike the Boom scenario, the Base Case does not require fast diffusion, broad gain-sharing and unconstrained capital deepening all at once — it only requires diffusion to proceed at something like the historical GPT pace, under the market structure we already observe, which is the lowest-assumption, highest-prior scenario of the four.
The Base Case is not a comfortable answer for anyone selling AI transformation at Goldman-scale multiples, nor for anyone hoping AI will single-handedly resolve a decade of weak productivity growth. It is, on the weight of two independently-derived institutional estimates and a wide base of firm-level field evidence, the most probable single outcome — which is exactly why the reframe in the closing section of this report matters: a Base Case aggregate can still conceal a Scenario 4 dynamic underneath it, region by region and sector by sector.
Scenario 3 — Stagnation / The Productivity Paradox Redux (~15–20% likelihood)
The stagnation case is not merely “the Base Case, but slower.” It is a structurally distinct claim: that diffusion stalls hard enough, or the investment boom decouples far enough from realized value, that AI’s net contribution to measured growth over the coming decade is close to zero — with a meaningful tail risk of an outright investment bust dragging growth briefly negative.
Erik Brynjolfsson, Daniel Rock and Chad Syverson’s NBER working paper “Artificial Intelligence and the Modern Productivity Paradox: A Clash of Expectations and Statistics” is the intellectual anchor here, and it is worth reading closely because it is not a pessimistic paper about AI’s ultimate potential — it is a paper about why potential and measured statistics can diverge for a long time. The authors open bluntly: “measured productivity growth has declined by half over the past decade, and real income has stagnated since the late 1990s for a majority of Americans” even as AI systems match or surpass human performance in more and more domains. They offer four explanations for the paradox — false hopes (the technology is simply less transformative than believed), mismeasurement (statistics fail to capture real gains), redistribution (private gains that are zero-sum at the aggregate level, e.g. one firm’s AI-driven market-share gain is another’s loss), and implementation lag — and conclude, after weighing the evidence, that implementation lags have likely been the biggest contributor to the paradox: the most impressive AI capabilities have not yet diffused widely, and like every prior general-purpose technology, their full effect will not be realized until waves of complementary innovation — organisational redesign, new skills, new business processes — catch up, a process the paper models as a form of unmeasured intangible capital investment.
The companion Productivity J-Curve paper formalises the mechanism precisely: measured productivity should be expected to dip before it rises, because firms are spending real resources on GPT-complementary intangible capital that standard national accounts do not capitalise, so the investment shows up as a cost with no offsetting asset on the books — a mechanically depressing effect on measured TFP that has nothing to do with AI’s true productive potential and everything to do with accounting convention. This is the same phenomenon the OECD’s own Annex A.2 addresses under the heading “Baumol’s growth disease” — a decomposition of how factor reallocation and relative-price changes can drag on aggregate productivity growth even while sector-level productivity is genuinely rising.
Ten supporting points for treating this as a real, not merely academic, near-term risk:
The theoretical ceiling on automation-driven growth is a Baumol constraint, not a technology constraint. Philippe Aghion, Benjamin F. Jones and Charles I. Jones’s NBER paper “Artificial Intelligence and Economic Growth” models AI as the latest wave of a 200-year automation process and finds that growth may ultimately be constrained not by what AI is good at but by what remains essential and yet hard to improve — the irreplaceable-task bottleneck that is the AI-era restatement of Baumol’s cost disease. If a large share of aggregate value-added sits in tasks AI cannot yet touch (skilled trades, elder care, complex judgment, physical world interaction), automating everything else at zero cost still caps aggregate growth at a rate set by the stubborn residual.
The AI capex boom is already showing the structural signature of a contest-theory overbuild, not an efficient one. The BIS’s “AI Investment Race” model — over-investment calibrated at 1.5–3x the efficient level, financed through debt and circular equity ties between hyperscalers, chipmakers and model labs, with cascading network exposure — is not a hypothetical; it is a description of financing structures already observed in balance-sheet and deal data as of the paper’s July 2026 publication.
NBER’s own investment-flow analysis is used to bound, not inflate, plausible cumulative GDP effects. “What Investment Data Implies about the AI Transition” (NBER Working Paper 35290) explicitly uses observed AI infrastructure investment to discipline what cumulative GDP effect is consistent with the capital actually being deployed — a methodological choice that, by construction, produces more conservative implied growth than a use-case-based forecast like McKinsey’s or PwC’s.
Energy is a hard, not soft, constraint on how fast the capex can even be converted into usable compute. The IEA’s “Energy and AI” special report and RAND’s “AI’s Power Requirements Under Exponential Growth” both model grid capacity — not chip supply — as a binding near-term bottleneck: transformers, transmission permitting and generation buildout all move on multi-year timelines that do not compress just because model capability is advancing on a compute-doubling cycle Epoch AI measures at roughly every six months in its “Rising Costs of Training Frontier AI Models.”
The historical base rate says stalling for a decade or more inside a multi-decade diffusion curve is the norm, not the exception. Paul David’s electricity study documents exactly this pattern — a long trough of disappointing productivity numbers between invention and full economic realisation — and Bresnahan and Trajtenberg’s original NBER theoretical treatment of general-purpose technologies, “Engines of Growth?”, explains why: GPTs require complementary innovation in every using-sector before their productivity potential is unlocked, and that complementary innovation is itself constrained by organisational and human-capital adjustment costs that do not move at software speed.
Mismeasurement can cut either way, and the false-hopes channel cannot be dismissed. Brynjolfsson, Rock and Syverson’s own taxonomy keeps “false hopes” — that generative AI’s real economic contribution is simply smaller than current enthusiasm implies, once hype-driven capital allocation is stripped out — as one of the four live explanations, and do not claim to have falsified it, only to have found implementation lag the larger of the four in their reading of the evidence.
The redistribution channel means some of the measured stagnation could be real even as reported corporate AI adoption rises. If AI’s gains are substantially redistributive — one firm’s win is a rival’s loss, netting to roughly zero at the aggregate level — then rising firm-level AI adoption metrics (the kind reported in surveys like the WEF’s “Future of Jobs Report 2025”) are compatible with flat aggregate productivity, exactly the paradox Brynjolfsson, Rock and Syverson set out to explain.
A financial-stability tail risk is explicit in the library, not merely implied. The BIS investment-race paper’s network analysis shows that stress in one AI-buildout firm could cascade to others through chains of financial exposure — meaning the downside of this scenario is not simply “slower growth than hoped” but includes a discrete probability of an investment-bust event that subtracts from measured growth for a period, echoing the dot-com capex cycle but with debt and circular financing structures the BIS paper flags as a distinct amplifying mechanism.
The OECD’s own scenario modelling treats slow-diffusion paths as a first-order sensitivity, not an edge case. “Miracle or Myth?” explicitly models the sensitivity of its central 0.25–0.6 percentage-point estimate to alternative assumptions about sectoral gains, demand response and reallocation frictions (its Figure 11), meaning the paper’s own central estimate already sits inside a distribution whose lower tail overlaps meaningfully with a stagnation outcome.
The gap between AI capability benchmarks and AI economic diffusion is now a measured, tracked quantity, and it is widening, not narrowing — Stanford HAI’s “2026 AI Index Report” is the annual instrument the library uses to track compute, investment and model-performance trends, and the persistence of a capability-diffusion gap year over year is itself evidence against the fast-diffusion assumption the Boom scenario requires and in favour of the multi-year lag this scenario describes.
ENSI Foresight Division prices Stagnation at roughly 15–20% — meaningfully more likely than the Boom scenario, because it requires only one thing to go wrong (diffusion friction, or an investment overbuild correcting) rather than three things to go right, but still a minority outcome because the weight of the firm-level field evidence in Angle 15 (the call-centre, developer and robot-adoption studies) shows AI is already delivering some measurable productivity gain at the point of deployment — the stagnation case therefore requires that gain to fail to aggregate up, not that it fails to exist at the micro level, which is a narrower and less probable claim than pure technological disappointment would be.
Scenario 4 — Demand-Driven Shrinkage (~10–15% likelihood)
This is the scenario a policymaker is least likely to have internalised, because it inverts the intuition built by every supply-side headline number in this report. Goldman’s 7%, PwC’s $15.7 trillion, Acemoglu’s 0.66%, the OECD’s 0.25–0.6 percentage points — every one of these is, at root, a statement about potential output: what the economy is capable of producing once AI’s productivity effect is fully realized. None of them is a full macroeconomic model that closes with a demand side and asks whether anyone has the income to buy what the supply side now can produce. Scenario 4 is what happens when you ask that second question and the answer is no.
The mechanism, stated precisely, is this: automation raises output-per-worker (labour productivity rises; potential GDP rises) but simultaneously reduces the wage income of displaced or wage-suppressed workers. If the Acemoglu-Restrepo reinstatement effect — new tasks created for displaced labour — and government redistribution together fail to replace that lost labour income fast enough, aggregate demand falls. Because realized GDP in the short-to-medium run is demand-constrained, not supply-constrained, realized GDP can fall even as output-per-worker (and therefore measured “productivity”) rises. This is not a contradiction; it is what happens when a supply-side productivity shock is fed into a demand-side model rather than assumed away. The same underlying technological event that reads as unambiguous “growth” in Goldman’s task-exposure framework or Acemoglu’s Hulten’s-theorem bound can read as contraction the moment the model is closed with a Keynesian or New Keynesian demand block instead of assumed at full employment and full income pass-through.
The BIS’s “The Impact of Artificial Intelligence on Output and Inflation” (Aldasoro, Doerr, Gambacorta and Rees, BIS Working Paper 1179) makes the pivot condition explicit and precise. Modelling AI as a permanent, sector-differentiated productivity shock inside a macroeconomic multi-sector model, the authors find that the entire character of the response — expansionary or contractionary, inflationary or disinflationary — hinges on whether households and firms anticipate the future productivity gain. In their own words: “if they do not anticipate higher future productivity, AI adoption is initially disinflationary” and only gradually becomes moderately inflationary through general-equilibrium demand effects; “in contrast, when households and firms anticipate higher future productivity, inflation rises immediately” as spending is pulled forward against expected future income. This is precisely the fork between Scenario 2 (anticipated, gain-sharing, demand recovers) and Scenario 4 (unanticipated or undelivered, demand lags, output can contract before it expands) — and which fork an economy lands on is an empirical, observable question about expectations formation and income distribution, not a technological one.
Anton Korinek and Joseph Stiglitz’s NBER working paper “Artificial Intelligence and Its Implications for Income Distribution and Unemployment” builds the taxonomy this scenario needs. They identify the two main channels through which AI-driven inequality operates — the surplus captured by innovators, and the redistribution that occurs through factor-price changes (i.e., wages falling as labour’s bargaining position weakens relative to capital) — and they formalise two distinct channels of technological unemployment: an efficiency-wage channel, where firms find they can pay less and still retain adequate labour once AI substitutes for worker leverage, and a transitional-phenomenon channel, where unemployment is a temporary but potentially long-lasting feature of the adjustment path rather than a new steady state. Crucially, Korinek and Stiglitz show that even in the theoretically favourable case where AI could deliver a Pareto improvement — where redistribution could in principle make everyone better off — that outcome requires non-distortionary taxation and active redistribution to actually occur. The Pareto-improvement result is conditional on policy action, not automatic; absent it, their model produces exactly the winners-and-losers dynamic that defines this scenario.
This is not a purely theoretical construction bolted onto abstract macro models. It has a real, close empirical precedent, and the library’s demand-side and market-structure angles converge on why it is plausible rather than merely logically possible:
Secular stagnation was already underway before AI, which lowers the bar for a demand-side shock to bite. Gauti Eggertsson, Neil Mehrotra and Lawrence Summers’ “Secular Stagnation in the Open Economy” formalises how a persistent demand shortfall can trap an economy in a low-growth equilibrium; Łukasz Rachel and Lawrence Summers’ “On Secular Stagnation in the Industrialized World” documents that neutral real interest rates across the industrialized bloc had already fallen by at least 300 basis points over the preceding generation on standard estimates — and by “as much as 700 basis points since the 1970s” on their private-sector-neutral-rate measure — reflecting chronic weakness in demand relative to desired saving, largely independent of AI. An economy already leaning toward demand insufficiency is an economy with less slack to absorb a labour-income shock without realized output actually falling.
The mechanism by which gains fail to reach households as wages is not hypothetical — it is the same market-power evidence documented in Scenario 1’s rebuttal. De Loecker and Eeckhout’s finding that average US markups have risen from roughly 18% to 67% above marginal cost since 1980, alongside their explicit finding that this coincides with a falling labour share, falling low-skill wages, and slower aggregate output growth, is the empirical channel through which “AI raises productivity but wages don’t rise proportionally” stops being a theoretical possibility and becomes a continuation of an already-observed thirty-year trend.
The scale of exposure is large enough for the demand channel to matter in aggregate, not just at the margin. The ILO’s “Generative AI and Jobs: A Refined Global Index of Occupational Exposure” and Goldman’s own two-thirds-of-jobs exposure estimate both describe an exposed labour share large enough that even a partial failure of the reinstatement effect would move aggregate labour income by a materially larger amount than any plausible near-term productivity offset.
Reallocation, empirically, takes far longer than models with instantaneous labour-market clearing assume — the China Shock is the closest real-world analogue in scale and mechanism. Autor, Dorn and Hanson’s “The China Shock” is explicit: “adjustment in local labor markets is remarkably slow, with wages and labor-force participation rates remaining depressed and unemployment rates remaining elevated for at least a full decade” after the shock begins, and “offsetting employment gains in other industries... have yet to materialize” even a decade-plus after the initial exposure. If an AI-driven demand/employment shock displaces workers at even a fraction of the geographic concentration the China Shock exhibited, the Autor-Dorn-Hanson evidence says the reinstatement effect Acemoglu’s Base Case model relies on cannot be assumed to arrive within any policy-relevant timeframe.
The Autor-Salomons finding that own-industry productivity gains reduce own-industry employment is the microfoundation for why aggregate demand, not just aggregate supply, needs to be modelled explicitly. Their ECB conference paper “Does Productivity Growth Threaten Employment?” finds industry-level productivity gains do reduce employment within the gaining industry, even where economy-wide effects vary — exactly the granular mechanism Scenario 4 needs to be more than a stylised macro story.
The direction of the demand response is not fixed by the technology — it is fixed by policy and expectations, which is itself the actionable finding. The BIS output-and-inflation paper’s central result — that the entire sign of the near-term response depends on anticipation — means Scenario 4 is not a prediction that shrinkage will happen; it is a precise statement of the condition under which it happens (unanticipated or undelivered gains, weak redistribution, concentrated market structure), which is exactly the kind of condition a government can monitor and intervene against.
The IMF’s own productivity-headwinds work shows how weak demand and balance-sheet damage produce lasting hysteresis, not a one-off dip. “Gone with the Headwinds: Global Productivity” documents how demand weakness and balance-sheet damage can scar potential output itself over time — meaning a Scenario-4-style demand shortfall is not necessarily self-correcting even after the initial adjustment period, if it damages investment and human capital accumulation along the way.
This scenario does not require universal wage stagnation — it requires the reinstatement-and-redistribution race to lose to the displacement-and-concentration race in enough of the economy, for long enough, to show up in realized output, which given the evidence above (thirty years of rising markups, a decade-plus of trade-shock evidence on reallocation speed, and an explicit BIS finding that the outcome hinges on anticipated income) is a materially more plausible bar to clear than “AI causes a literal, economy-wide, permanent contraction.”
ENSI Foresight Division prices outright, sustained national demand-driven shrinkage at roughly 10–15% — a real, non-trivial probability, but a minority one, because governments retain fiscal and monetary tools (the BIS output-and-inflation paper’s own point about anticipation being manageable through policy communication and demand management) that make a full national contraction an avoidable rather than an inevitable outcome. But — and this is the hinge on which the rest of this report turns — pricing the national aggregate version of this scenario at 10–15% radically understates how often the underlying mechanism fires at sub-national scale, which is exactly the gap the Bifurcation reframe below is built to close.
One paragraph on where the agentic layer fits — and why it belongs to Report 3
Everything above treats AI as a tool that humans deploy inside existing firms, tasks and labour markets. A qualitatively different mechanism is emerging alongside it: AI agents that transact, negotiate and compete as autonomous economic actors rather than simply assisting a human worker. Gillian Hadfield and Andrew Koh’s “An Economy of AI Agents,” prepared for the NBER Handbook on the Economics of Transformative AI, surveys how agents capable of planning and executing complex tasks over long time horizons with little human oversight could reorganise markets and institutions in ways the task-based automation literature was not built to anticipate. The University of Cambridge’s “When AI Agents Compete for Jobs” offers an early empirical glimpse of the dynamics this could produce: in a simulated AI labour market, agents equipped with metacognition, competitive awareness and long-horizon strategic planning achieve roughly 1.5 times the market share of agents using standard prompting approaches, and the simulation exhibits price-deflation and market-concentration dynamics as capable agents out-compete less capable ones at machine speed. Whichever of the four scenarios above turns out to dominate, an agentic layer would not create a fifth outcome so much as accelerate the arrival of whichever one is already winning — compressing the Boom scenario’s diffusion timeline, deepening the Stagnation scenario’s overbuild, or intensifying the Demand-Shrinkage scenario’s concentration mechanics. This is deliberately not developed further here: it is the entire subject of Report 3 in this series, and flagging the hand-off rather than pre-empting it is the correct discipline for Report 1.
The Bifurcation — the reframe, stated directly
Here is the argument this report has been building toward: the sharpest, most underpriced risk in the AI-growth debate is not which of the four scenarios above the global economy lands in. It is that the question is being asked at the wrong level of aggregation.
Every headline forecast surveyed in this report — Goldman’s 7%, PwC’s $15.7 trillion, Acemoglu’s 0.66%, the OECD’s 0.25–0.6 percentage points, the IMF’s advanced-versus-low-income growth differential — is a national or global aggregate, built from a supply-side, task-exposure or general-equilibrium model that does not fully endogenize the demand-side feedback loop described in Scenario 4. This is not a criticism of the methodology on its own terms: Acemoglu’s paper is explicitly a supply-side bound; the OECD’s is a productivity-focused general-equilibrium exercise; Goldman’s is a task-exposure extrapolation. Each is doing what it says it does. The problem is what happens when a policymaker reads the aggregate number as if it were a description of what will happen everywhere inside the country, uniformly, when the models were never built to make that claim in the first place.
Consider the arithmetic directly. The IMF’s “Mind the Gap” finds AI’s estimated growth effect could be more than double in advanced versus low-income economies — that is already a statement that the “0.4pp of global TFP” aggregate is a blend of very different sub-experiences. Push the same disaggregation down one more level, from country to region and sector within a country, and the same logic applies with more force, not less, because sub-national labour markets are less mobile and less fiscally insulated than national economies are from each other. The BIS’s cross-country study explicitly finds AI’s growth effect concentrated in advanced economies at the country level; nothing in the model architecture prevents an equivalent concentration within an advanced economy, between its AI-exposed financial and professional-services metros and its AI-exposed-but-poorly-reinstated manufacturing and back-office regions. A national GDP figure that averages a compute-cluster metro seeing a Goldman-style boom with three deindustrialising service corridors seeing a Korinek-Stiglitz-style demand contraction can report as an entirely unremarkable Base Case number — Acemoglu’s 0.5%, the OECD’s 0.4 percentage points — while concealing exactly the outcome this report’s readers most need to see coming.
The China Shock is the closest real-world precedent in the library for how long-lasting and geographically concentrated this kind of shock can be, and it is worth stating why it is the right analogue rather than a loose metaphor. Trade exposure to China, like AI exposure, was not evenly distributed — it hit specific industries (furniture, textiles, electronics assembly) concentrated in specific commuting zones. Autor, Dorn and Hanson’s evidence that those commuting zones saw depressed wages, depressed labour-force participation and elevated unemployment for at least a full decade, with offsetting employment gains in other industries that “have yet to materialize” even years later, is precisely the pattern a geographically- and sectorally-concentrated AI exposure shock would be expected to produce — and precisely the pattern that a national GDP or national employment aggregate is constructed to average away. The US economy as a whole absorbed the China Shock without a recession attributable to trade alone; the counties and workers directly exposed experienced something that looked, to them, indistinguishable from Scenario 4, for a decade, inside an aggregate reading Scenario 2.
This is why ENSI Foresight Division treats the Bifurcation not as a fifth item in the probability table but as a modifier that sits on top of all four scenarios simultaneously. Whichever aggregate outcome wins — even the Boom scenario — the distributional mechanics documented in the market-structure and labour-reallocation angles of this library (rising markups, superstar-firm concentration, decade-plus reallocation lags, more-than-double advanced-versus-low-income growth differentials) do not disappear inside a good aggregate number. They are, if anything, easier to miss inside a good number, because a rising national GDP figure is exactly the kind of signal that makes a finance ministry stop asking where the gains are landing.
What to watch: leading indicators for policymakers
A state cannot wait for the ten-year TFP number to resolve before deciding how to prepare. The following indicators, each drawn from a mechanism documented above, are observable well before the headline growth print confirms which scenario is materializing — and, read together, they are built specifically to catch the Bifurcation the aggregate number would miss.
Track AI-exposed employment and wage data at the regional and occupational level, not only the national level. The ILO’s occupational-exposure index and the OECD’s sectoral-exposure mapping (used to build the “Miracle or Myth?” general-equilibrium model) already exist as instruments; the discipline is to run them quarterly against regional labour-force data the way the Autor-Dorn-Hanson China Shock methodology tracked commuting-zone-level trade exposure, rather than waiting for a national unemployment print to move.
Watch the gap between AI capital expenditure and realized enterprise value creation. The BIS investment-race paper’s over-investment ratio (1.5–3x efficient level) and the NBER’s investment-flow-implied GDP bound are both designed to be recomputed as new balance-sheet and deal data arrive; a widening gap between announced AI capex and delivered productivity gain is the earliest observable signal of the Stagnation scenario’s bust tail, and it is visible years before an aggregate recession would be.
Monitor markup and labour-share trends by sector, specifically in AI-adjacent industries. If cloud, compute and frontier-model markups continue the De Loecker-Eeckhout trajectory (18% to 67% above marginal cost since 1980, concentrated in a rising share of high-markup firms) rather than compressing as AI capability commoditises, that is a direct, sector-level signal that Scenario 4’s rent-capture mechanism — not broad gain-sharing — is the one operating.
Track household and firm expectations of future income directly, not just realized income. The BIS output-and-inflation paper’s central finding — that the entire sign of the demand response depends on anticipation — means survey-based expectations data (of the kind central banks already collect for inflation expectations) can be repurposed as an early-warning instrument for whether an economy is on the Scenario 2 or Scenario 4 path, months or years before realized GDP data would reveal it.
Watch reallocation speed against the China Shock’s decade-long benchmark, not against an assumed frictionless model. If displaced-worker reemployment rates, wage recovery timelines and labour-force-participation trends in AI-exposed regions are tracking faster than the China Shock’s documented decade-plus adjustment period, that is meaningful evidence the reinstatement effect is working in something closer to real time — a genuinely different, more optimistic signal than if they are tracking the same or slower.
Watch neutral real interest rates and secular-stagnation indicators as a macro backdrop, not a separate story. Rachel and Summers’ finding of a 300–700 basis point decline in neutral real rates over the preceding half-century describes the demand-side terrain AI’s productivity shock is landing on; a further decline, rather than stabilisation or reversal, would indicate the economy has even less capacity to absorb a labour-income shock without realized output actually falling — raising, not lowering, the probability mass on Scenario 4.
Track energy and grid-capacity data as a hard constraint on diffusion speed, independent of model-capability announcements. The IEA and RAND data on data-centre power demand and grid buildout timelines are a genuinely exogenous check on how fast the Boom or Base Case scenarios can physically proceed, regardless of how fast frontier-model capability itself advances — a useful discipline against extrapolating growth forecasts purely from compute-scaling trends.
Watch for the emergence of agentic, autonomous AI transaction volume as an accelerant signal, not a fifth scenario. Early indicators — the kind Anthropic’s own Economic Index Report already tracks in real Claude usage data on task delegation and autonomy — are worth monitoring as a leading sign of how fast whichever scenario is already winning will arrive, which is the hand-off this report makes explicitly to Report 3.
The bifurcation, restated
The honest answer to “will AI grow or shrink the economy” is: probably grow it, modestly, unevenly, over a longer timeline than the enthusiasts assume and a shorter one than the historical base rate would predict for a typical general-purpose technology — while concealing, inside that same modest positive aggregate, regional and sectoral episodes that will look and feel like outright shrinkage to the people living through them, for years, exactly as the China Shock did. The four scenarios in this report are not a menu from which the world will choose one. The Base Case is where ENSI Foresight Division places the largest single share of probability, at roughly 45–50%, anchored to the convergence of Acemoglu’s task-based bound and the OECD’s independent general-equilibrium estimate. But the number that should worry a policymaker most is not the probability attached to any of the four rows in the summary table above — it is the near-certainty that whichever row wins, the Bifurcation modifier is already operating underneath it, and that the standard forecasting toolkit used by Goldman, McKinsey, the OECD and the IMF is not built to show a finance minister where. A state that only watches its national GDP print will find out it was living through Scenario 4 at the regional level roughly a decade after the fact — which is exactly how long it took to fully recognise the China Shock for what it was. The leading indicators above exist so that this report’s readers do not have to wait that long again.




