Five Futures — Will AI Grow or Shrink the Economy?

September 5, 2026
blog image

Built on a library of 105 primary documents spanning the NBER, IMF, OECD, BIS, Federal Reserve System, World Bank and the major AI labs and forecasters — ENSI Foresight Division.


The forecast that hides the thing it should be measuring

Every institutional forecast of AI’s economic effect — Goldman Sachs’ 7%, PwC’s $15.7 trillion, the OECD’s 0.25–0.6 percentage points, the IMF’s country-by-country growth differentials — answers a narrower question than the one policymakers think they are asking. Each is, in its own methodology, a model of potential output: what the economy could produce if the productivity gain shows up and is spent. None of the headline numbers that dominate boardroom slides and finance-ministry briefings is a model of realized, demand-constrained GDP — the number that determines whether tax receipts rise, whether unemployment claims fall, and whether a government gets re-elected. That gap between potential and realized output is not a technicality. It is the single most consequential and most underpriced risk in the entire AI-and-growth literature, and this report exists to make it visible.

The reframe this report argues for is simple to state and easy to miss: the modal, mainstream forecast — modest, positive, uneven aggregate growth — is entirely compatible with regions, sectors and income deciles inside that aggregate experiencing outright contraction, for the same reason and at the same time. A national GDP print of +0.4% is arithmetically consistent with call-centre towns, back-office cities and mid-skill service corridors losing income for a decade, exactly as Pittsburgh, Greensboro and the furniture counties of North Carolina did during the “China Shock” — not because the aggregate model was wrong, but because aggregate models are built to net out precisely the geography and distribution that determines whether a shock feels like growth or like shrinkage to the people living through it. Aggregation is not measurement error. It is a modelling choice, and every one of the headline AI-growth forecasts in this library makes it.

This matters because the entities that most need an accurate scenario map — national treasuries, central banks, regional development agencies, sovereign wealth funds, multilateral lenders — are not making a bet on a single global GDP number. They are making a portfolio of decisions about fiscal buffers, retraining budgets, monetary stance, regional transfers and industrial policy, each of which depends on knowing not just whether AI grows the pie but whose slice moves and how fast the redistribution happens relative to the political cycle that has to absorb it. A state that plans against Goldman Sachs’ 7% uplift and gets the OECD’s 0.4 percentage points instead has made a forecasting error. A state that plans against any aggregate figure at all, while a fifth of its regions or a third of its labour force experiences the Korinek-Stiglitz demand-shrinkage mechanism described in Scenario 4 below, has made a category error — and it is the more dangerous of the two, because the national dashboard will keep reading “growth” the entire time.

The five-scenario framework that follows is deliberately not a spectrum from “big growth” to “big shrinkage” with the truth somewhere in the middle. It is four distinct causal stories — a productivity boom, a modest base case, a stagnation trap, and a demand-driven contraction — each anchored to a different published model with a different methodology and a different set of assumptions about diffusion speed, gain-sharing and constraint-bindingness, plus a fifth argument that cuts across all four: that the geographic and sectoral variance inside whichever aggregate number wins is where the real policy risk lives. Readers should leave this report able to name, for their own jurisdiction, which leading indicators would move them from one scenario to another — and understanding that the question “will AI grow or shrink the economy” is underspecified until you also ask “for whom, and measured how.”

One discipline runs through every section below: history’s base rate for how long general-purpose technologies take to resolve into measured productivity. Paul David’s canonical study, “The Dynamo and the Computer,” published in the American Economic Review, found that electrification took roughly forty years to show up in US productivity statistics after the technology existed — because factories had to be rebuilt around unit-drive motors rather than retrofitted around a single central power source, and that reorganisation, not the invention, was the rate-limiting step. Nicholas Crafts’ growth-accounting study for the LSE, “Steam as a General Purpose Technology,” pushes the base rate further: steam power took close to a century to lift UK aggregate productivity growth measurably. Every probability estimate in this report should be read against that discipline. A scenario “resolving” by 2030 or even 2035 would be fast by the standard of the only two general-purpose technologies we have full-century data on.

The five futures, at a glance

  • The Productivity Boom (~15% likelihood). Fast diffusion, broadly shared gains, no binding physical constraint — the world Goldman Sachs and PwC model. Attractive, well-funded, and the least likely of the four core scenarios because it requires three separate things the historical and structural evidence argues against simultaneously.

  • The Base Case — Modest, Uneven Growth (~45–50% likelihood). The mainstream institutional consensus: Acemoglu’s task-based ceiling, the OECD’s general-equilibrium range, the IMF and BIS’s cross-country findings. Growth that is real, arrives slowly, and is unevenly distributed by construction — the most probable single outcome and the one this report spends the most time defending.

  • Stagnation / The Productivity Paradox Redux (~15–20% likelihood). Diffusion stalls the way it stalled for computers in the 1980s and 1990s; the AI capex boom outruns realized value the way the BIS’s contest-theory model predicts; Baumol’s cost disease caps what automation alone can deliver.

  • Demand-Driven Shrinkage (~10–15% likelihood). The scenario readers most need explained because it is the least intuitive: potential output rises even as realized GDP falls, because the same productivity shock that looks expansionary in a supply-side model becomes contractionary once you close the model with a demand side that depends on wage income households don’t have.

  • The Bifurcation (the report’s central reframe, not a fifth probability bucket). A “modest growth” aggregate (Scenario 2) can mechanically conceal regional and sectoral versions of Scenario 4 playing out underneath it — because the published aggregate models are not built to catch that. This is the sharpest, most underpriced risk in the whole literature, and it is why “which scenario wins nationally” is the wrong question for any policymaker to be asking alone.


Scenario 1 — The Productivity Boom (~15% likelihood)

The bull case is not a fringe position; it is argued by two of the most-cited institutions in applied macro forecasting, and its arithmetic is worth taking seriously before it is discounted. Goldman Sachs’ Joseph Briggs and Devesh Kodnani, in “The Potentially Large Effects of Artificial Intelligence on Economic Growth,” estimate that roughly two-thirds of current US and European jobs are exposed to some degree of AI automation, that generative AI could substitute for up to one-quarter of current work tasks, and that extrapolated globally this is equivalent to exposing the labour input of some 300 million full-time jobs to automation. Run through their model, that yields just under 1.5 percentage points of additional annual US labour-productivity growth over a ten-year period following widespread adoption, and an eventual 7% increase in global GDP — about $7 trillion in their 2023 pricing. PwC’s “Sizing the Prize” is, if anything, the more dramatic of the two headline forecasts in this library: $15.7 trillion of additional global GDP by 2030, equivalent to global output being up to 14% higher than it would otherwise be, split by PwC’s own accounting into productivity effects (businesses automating processes and augmenting labour, which the firm says account for over 55% of total GDP gains between 2017 and 2030) and a second, growing consumption-side effect as AI-enhanced products and personalisation drive additional demand — a channel PwC says will account for 58% of the GDP gain realized in 2030 specifically, i.e. the mix shifts from productivity-led to demand-led as the diffusion matures. PwC further disaggregates by geography and sector: China stands to see the largest proportional boost (up to 26% of GDP), North America the second largest (14%), and retail, financial services and healthcare the sectors with the greatest combined productivity-and-product-enhancement potential.

Why does ENSI Foresight Division treat this as the low-probability tail rather than the central case, despite the credibility of the institutions behind it? Three reasons, each traceable to a different angle of this library, and each a condition the Boom scenario requires to hold simultaneously — which is precisely what makes their joint probability low even when each is individually plausible.

  • It requires diffusion speed that the historical record argues against. Paul David’s electricity study and Nicholas Crafts’ steam study — the two general-purpose technologies with a full-century data record — took 40 and roughly 100 years respectively to show up in aggregate productivity, because the complementary reorganisation of firms, not the invention itself, was the binding constraint. The Federal Reserve Board’s 2025 working paper “Generative AI at the Crossroads: Light Bulb, Dynamo, or Microscope?” tests generative AI explicitly against these historical diffusion analogues and finds the honest answer is not yet knowable — the paper’s title is itself an admission that AI could be any of the three, with very different growth implications.

  • It requires broad gain-sharing that the market-structure evidence directly contradicts. Jan De Loecker and Jan Eeckhout’s NBER working paper “The Rise of Market Power and the Macroeconomic Implications” documents that average US markups rose from roughly 18% above marginal cost in 1980 to 67% today, concentrated in an increase in high-markup firms rather than a broad-based shift — and they link this rise causally to falling labour share, falling low-skill wages, and a slowdown in aggregate output. Autor, Dorn, Katz, Patterson and Van Reenen’s companion NBER paper, “The Fall of the Labor Share and the Rise of Superstar Firms,” finds the same winner-take-most dynamic in industry concentration data. If AI’s productivity gains flow disproportionately to the handful of firms that already control the compute, the cloud infrastructure and the frontier models — exactly the concentration the UK Competition and Markets Authority’s “AI Foundation Models” technical report and the FTC’s 6(b) study of cloud-AI partnerships were opened to investigate — then the Boom scenario’s assumption of broadly shared productivity gains breaks down at the first link in the chain.

  • It requires no binding compute or energy constraint, which Angle 04’s evidence says is not a safe assumption. The IEA’s “Energy and AI” special report models data-centre electricity demand as a potential bottleneck on how fast AI compute can scale at all — grid capacity, transformer lead times and permitting are not software problems that can be diffused at ChatGPT’s adoption speed. The BIS’s “The AI Investment Race” (Phurichai Rungcharoenkitkul, July 2026) goes further and models the entire AI buildout as a winner-take-most contest in which competing firms rationally over-commit resources: calibrated to balance-sheet and deal data, the paper finds over-investment running at roughly 1.5 times the efficient level, rising to around 3 times where demand is less elastic — financed substantially through debt and circular equity ties between hyperscalers, model labs and chip suppliers, with a network-cascade risk if any one node disappoints on revenue.

None of this means the Boom scenario is impossible — Goldman and PwC’s methodologies are serious and their authors are not naïve about diffusion lags. It means the scenario requires fast diffusion, broad sharing and unconstrained capital deepening to all hold at once, against a historical base rate, a market-structure trend and a physical-infrastructure constraint that each independently argue the opposite. Compounding three low-conditional-probability requirements is why ENSI Foresight Division prices this scenario at roughly 15% — high enough to take seriously in capital allocation and infrastructure planning, low enough that a state should not build its ten-year fiscal plan on it arriving.

Scenario 2 — The Base Case: Modest, Uneven Growth (~45–50% likelihood)

This is the mainstream institutional consensus, and it deserves to be argued for on its own terms rather than treated as the residual “everyone else’s number.” Its intellectual anchor is Daron Acemoglu’s “The Simple Macroeconomics of AI,” prepared for Economic Policy and circulated as NBER Working Paper 32487. Acemoglu builds a task-based model in which AI’s macroeconomic effect is bounded by a version of Hulten’s theorem: aggregate productivity gains are given by the fraction of tasks AI actually touches multiplied by the average task-level cost saving. Using the best available exposure and productivity-improvement estimates, he finds the resulting TFP effect is “nontrivial but modest — no more than a 0.66% increase in total factor productivity over 10 years.” He then argues this may still be an overestimate, because current evidence is drawn disproportionately from easy-to-learn tasks, while much of AI’s future effect will have to come from hard-to-learn tasks with context-dependent judgment and no objective performance metric to train against — on that adjustment, his predicted 10-year TFP gain falls to under 0.53%. This is not a paper written to be contrarian for its own sake: Acemoglu explicitly engages Goldman’s 7% and McKinsey’s $17.1–25.6 trillion range in his introduction and argues the gap is explained by his more conservative estimate of which tasks are genuinely automatable versus merely exposed.

The OECD’s “Miracle or Myth? Assessing the Macroeconomic Productivity Gains from Artificial Intelligence” (OECD Artificial Intelligence Papers No. 29, November 2024) arrives independently at a strikingly similar order of magnitude through an entirely different method — a novel micro-to-macro multi-sector general-equilibrium model with input-output linkages, rather than Acemoglu’s task-exposure algebra. Its headline finding: annual aggregate TFP growth attributable to AI of 0.25–0.6 percentage points, equivalent to 0.4–0.9 percentage points of labour productivity growth over a ten-year horizon. That two independent methodologies — one a stylised task-based bound, one a full general-equilibrium simulation with sectoral input-output linkages — converge on the same order of magnitude, an order of magnitude roughly a fifth to a tenth the size of Goldman’s or PwC’s headline figures, is the single strongest piece of evidence in this library for treating the Base Case as the modal outcome rather than a competitor to the Boom scenario.

Ten to fifteen supporting reasons this is where ENSI Foresight Division places the largest probability mass:

  • It is where two independent, methodologically distinct institutional estimates converge, as above — Acemoglu’s task-exposure bound and the OECD’s general-equilibrium simulation, arrived at without coordination, land in the same 0.3–0.9 percentage-point-of-productivity-growth range.

  • It matches the observed pattern of real deployment, not just theory. The BIS’s “Artificial Intelligence and Growth in Advanced and Emerging Economies: Short-Run Impact,” a 56-economy, 16-industry cross-country study, finds a real but modest growth effect concentrated in advanced economies, consistent with a diffusion process still in its early innings rather than a step-change already realized.

  • It is consistent with firm-level field evidence showing real but bounded gains. Erik Brynjolfsson, Danielle Li and Lindsey Raymond’s “Generative AI at Work” studies 5,179 customer-support agents given access to a generative AI assistant and finds a 14% average productivity gain — economically significant, globally scalable in principle, but a 14% task-level gain is a different order of magnitude from a Goldman-style productivity boom, and the paper finds the gain is concentrated among novice and low-skilled workers (34% improvement) with minimal effect on already-experienced staff — a distributional pattern, not a uniform lift.

  • Software-development field experiments tell the same story of real, bounded, unevenly distributed gains. Microsoft Research’s three-firm RCT across 4,867 developers finds a 26% increase in task completion; the earlier GitHub Copilot field experiment finds developers completed a standardised task 55.8% faster. These are large individual-task effects that nonetheless net out, at the Acemoglu-style aggregate level, to a modest TFP contribution once weighted by the share of the economy actually composed of tasks this exposed.

  • The historical diffusion base rate argues for “modest and slow” over “large and fast.” The Productivity J-Curve literature (NBER Working Paper 25148) formalises why GPT adoption should be expected to show up first as a dip in measured productivity — as firms spend on complementary intangible investment that national accounts do not capture as capital — before any acceleration appears, which is exactly consistent with a base case that looks unremarkable for years before it looks real.

  • It is a growth effect real enough to matter for a finance ministry, wrong enough to disappoint an equity analyst pricing a Goldman-sized re-rating — which is itself diagnostic. Sustained modest TFP acceleration compounds: even Acemoglu’s more conservative 0.53% figure, sustained and extended past the initial ten-year window as diffusion continues per the historical base rate above, is not nothing over a multi-decade horizon.

  • It is unevenly distributed by construction, not by exception — the IMF’s “The Global Impact of AI: Mind the Gap” (WP/25/76) finds the estimated growth impact in advanced economies could be more than double that in low-income economies once sectoral exposure, technological preparedness and data/technology access are fed into a multi-region dynamic general-equilibrium model, with AI-driven productivity gains concentrated in the non-tradable sector large enough to disrupt the traditional exchange-rate-adjustment mechanism (an inverse Balassa-Samuelson effect, in the paper’s own framing).

  • It coexists with — rather than reverses — the concentration trend already documented in Angle 06. Nothing about a modest aggregate TFP gain requires that gain to be evenly shared; De Loecker and Eeckhout’s markup evidence and the superstar-firm literature describe a economy-wide trend already three decades in train, and the Base Case scenario simply assumes AI does not interrupt it, which is the more conservative and more defensible assumption than assuming AI reverses it.

  • It is what “Miracle or Myth?” itself frames as the honest middle — the OECD paper’s own Figure 1, comparing predicted macro-level productivity gains across the published studies, is explicitly built to show how much the estimates vary depending on methodology, and situates its own general-equilibrium estimate as the more disciplined, assumption-transparent number against which the use-case-based (McKinsey) and task-exposure (Goldman) estimates should be read as upper bounds rather than central forecasts.

  • It requires the fewest simultaneous strong assumptions. Unlike the Boom scenario, the Base Case does not require fast diffusion, broad gain-sharing and unconstrained capital deepening all at once — it only requires diffusion to proceed at something like the historical GPT pace, under the market structure we already observe, which is the lowest-assumption, highest-prior scenario of the four.

The Base Case is not a comfortable answer for anyone selling AI transformation at Goldman-scale multiples, nor for anyone hoping AI will single-handedly resolve a decade of weak productivity growth. It is, on the weight of two independently-derived institutional estimates and a wide base of firm-level field evidence, the most probable single outcome — which is exactly why the reframe in the closing section of this report matters: a Base Case aggregate can still conceal a Scenario 4 dynamic underneath it, region by region and sector by sector.

Scenario 3 — Stagnation / The Productivity Paradox Redux (~15–20% likelihood)

The stagnation case is not merely “the Base Case, but slower.” It is a structurally distinct claim: that diffusion stalls hard enough, or the investment boom decouples far enough from realized value, that AI’s net contribution to measured growth over the coming decade is close to zero — with a meaningful tail risk of an outright investment bust dragging growth briefly negative.

Erik Brynjolfsson, Daniel Rock and Chad Syverson’s NBER working paper “Artificial Intelligence and the Modern Productivity Paradox: A Clash of Expectations and Statistics” is the intellectual anchor here, and it is worth reading closely because it is not a pessimistic paper about AI’s ultimate potential — it is a paper about why potential and measured statistics can diverge for a long time. The authors open bluntly: “measured productivity growth has declined by half over the past decade, and real income has stagnated since the late 1990s for a majority of Americans” even as AI systems match or surpass human performance in more and more domains. They offer four explanations for the paradox — false hopes (the technology is simply less transformative than believed), mismeasurement (statistics fail to capture real gains), redistribution (private gains that are zero-sum at the aggregate level, e.g. one firm’s AI-driven market-share gain is another’s loss), and implementation lag — and conclude, after weighing the evidence, that implementation lags have likely been the biggest contributor to the paradox: the most impressive AI capabilities have not yet diffused widely, and like every prior general-purpose technology, their full effect will not be realized until waves of complementary innovation — organisational redesign, new skills, new business processes — catch up, a process the paper models as a form of unmeasured intangible capital investment.

The companion Productivity J-Curve paper formalises the mechanism precisely: measured productivity should be expected to dip before it rises, because firms are spending real resources on GPT-complementary intangible capital that standard national accounts do not capitalise, so the investment shows up as a cost with no offsetting asset on the books — a mechanically depressing effect on measured TFP that has nothing to do with AI’s true productive potential and everything to do with accounting convention. This is the same phenomenon the OECD’s own Annex A.2 addresses under the heading “Baumol’s growth disease” — a decomposition of how factor reallocation and relative-price changes can drag on aggregate productivity growth even while sector-level productivity is genuinely rising.

Ten supporting points for treating this as a real, not merely academic, near-term risk:

  • The theoretical ceiling on automation-driven growth is a Baumol constraint, not a technology constraint. Philippe Aghion, Benjamin F. Jones and Charles I. Jones’s NBER paper “Artificial Intelligence and Economic Growth” models AI as the latest wave of a 200-year automation process and finds that growth may ultimately be constrained not by what AI is good at but by what remains essential and yet hard to improve — the irreplaceable-task bottleneck that is the AI-era restatement of Baumol’s cost disease. If a large share of aggregate value-added sits in tasks AI cannot yet touch (skilled trades, elder care, complex judgment, physical world interaction), automating everything else at zero cost still caps aggregate growth at a rate set by the stubborn residual.

  • The AI capex boom is already showing the structural signature of a contest-theory overbuild, not an efficient one. The BIS’s “AI Investment Race” model — over-investment calibrated at 1.5–3x the efficient level, financed through debt and circular equity ties between hyperscalers, chipmakers and model labs, with cascading network exposure — is not a hypothetical; it is a description of financing structures already observed in balance-sheet and deal data as of the paper’s July 2026 publication.

  • NBER’s own investment-flow analysis is used to bound, not inflate, plausible cumulative GDP effects. “What Investment Data Implies about the AI Transition” (NBER Working Paper 35290) explicitly uses observed AI infrastructure investment to discipline what cumulative GDP effect is consistent with the capital actually being deployed — a methodological choice that, by construction, produces more conservative implied growth than a use-case-based forecast like McKinsey’s or PwC’s.

  • Energy is a hard, not soft, constraint on how fast the capex can even be converted into usable compute. The IEA’s “Energy and AI” special report and RAND’s “AI’s Power Requirements Under Exponential Growth” both model grid capacity — not chip supply — as a binding near-term bottleneck: transformers, transmission permitting and generation buildout all move on multi-year timelines that do not compress just because model capability is advancing on a compute-doubling cycle Epoch AI measures at roughly every six months in its “Rising Costs of Training Frontier AI Models.”

  • The historical base rate says stalling for a decade or more inside a multi-decade diffusion curve is the norm, not the exception. Paul David’s electricity study documents exactly this pattern — a long trough of disappointing productivity numbers between invention and full economic realisation — and Bresnahan and Trajtenberg’s original NBER theoretical treatment of general-purpose technologies, “Engines of Growth?”, explains why: GPTs require complementary innovation in every using-sector before their productivity potential is unlocked, and that complementary innovation is itself constrained by organisational and human-capital adjustment costs that do not move at software speed.

  • Mismeasurement can cut either way, and the false-hopes channel cannot be dismissed. Brynjolfsson, Rock and Syverson’s own taxonomy keeps “false hopes” — that generative AI’s real economic contribution is simply smaller than current enthusiasm implies, once hype-driven capital allocation is stripped out — as one of the four live explanations, and do not claim to have falsified it, only to have found implementation lag the larger of the four in their reading of the evidence.

  • The redistribution channel means some of the measured stagnation could be real even as reported corporate AI adoption rises. If AI’s gains are substantially redistributive — one firm’s win is a rival’s loss, netting to roughly zero at the aggregate level — then rising firm-level AI adoption metrics (the kind reported in surveys like the WEF’s “Future of Jobs Report 2025”) are compatible with flat aggregate productivity, exactly the paradox Brynjolfsson, Rock and Syverson set out to explain.

  • A financial-stability tail risk is explicit in the library, not merely implied. The BIS investment-race paper’s network analysis shows that stress in one AI-buildout firm could cascade to others through chains of financial exposure — meaning the downside of this scenario is not simply “slower growth than hoped” but includes a discrete probability of an investment-bust event that subtracts from measured growth for a period, echoing the dot-com capex cycle but with debt and circular financing structures the BIS paper flags as a distinct amplifying mechanism.

  • The OECD’s own scenario modelling treats slow-diffusion paths as a first-order sensitivity, not an edge case. “Miracle or Myth?” explicitly models the sensitivity of its central 0.25–0.6 percentage-point estimate to alternative assumptions about sectoral gains, demand response and reallocation frictions (its Figure 11), meaning the paper’s own central estimate already sits inside a distribution whose lower tail overlaps meaningfully with a stagnation outcome.

  • The gap between AI capability benchmarks and AI economic diffusion is now a measured, tracked quantity, and it is widening, not narrowing — Stanford HAI’s “2026 AI Index Report” is the annual instrument the library uses to track compute, investment and model-performance trends, and the persistence of a capability-diffusion gap year over year is itself evidence against the fast-diffusion assumption the Boom scenario requires and in favour of the multi-year lag this scenario describes.

ENSI Foresight Division prices Stagnation at roughly 15–20% — meaningfully more likely than the Boom scenario, because it requires only one thing to go wrong (diffusion friction, or an investment overbuild correcting) rather than three things to go right, but still a minority outcome because the weight of the firm-level field evidence in Angle 15 (the call-centre, developer and robot-adoption studies) shows AI is already delivering some measurable productivity gain at the point of deployment — the stagnation case therefore requires that gain to fail to aggregate up, not that it fails to exist at the micro level, which is a narrower and less probable claim than pure technological disappointment would be.

Scenario 4 — Demand-Driven Shrinkage (~10–15% likelihood)

This is the scenario a policymaker is least likely to have internalised, because it inverts the intuition built by every supply-side headline number in this report. Goldman’s 7%, PwC’s $15.7 trillion, Acemoglu’s 0.66%, the OECD’s 0.25–0.6 percentage points — every one of these is, at root, a statement about potential output: what the economy is capable of producing once AI’s productivity effect is fully realized. None of them is a full macroeconomic model that closes with a demand side and asks whether anyone has the income to buy what the supply side now can produce. Scenario 4 is what happens when you ask that second question and the answer is no.