Skip to content
24 min read

The bottleneck keeps moving. Right now it's electricity. · Strategist edition

In September 2024, Microsoft signed a deal to restart Three Mile Island, a nuclear plant shuttered since 2019. The press called it an energy story. It was an infrastructure story: Microsoft needed power for data centers and could not get it from the grid fast enough.

In September 2024, Microsoft signed a deal to restart Three Mile Island, a nuclear plant shuttered since 2019. The press covered it as energy policy. It was an infrastructure story: Microsoft needed power for data centers on a timeline that the grid could not match.

Four months later, Meta announced it was pursuing 6.6 gigawatts of nuclear generation. Those plants do not deliver until after 2030. The infrastructure they are meant to support is being built now.

The distance between when AI infrastructure needs power and when the grid can supply it is the constraint defining this industry in 2026, and probably through 2028. Not chips. Not capital. The wire from the substation to the building.

The panel's view

Scenario Probability
Grid Claims the Cascade: power/interconnection dominant by 2028 35%
Multi-Front Deadlock: HBM friction + grid bind simultaneously 25%
Efficiency Dampens the Cascade: AI demand decelerates before 2028 20%
Contingency Reshuffles the Stack: architecture or geopolitical pivot 20%

No calls have resolved yet. This is the panel's opening position.

Three of four frameworks converge on grid as the dominant binding constraint by 2027-28. The fourth dissents only on severity, not direction. Combined, the two grid-dominant scenarios carry a 60% probability.

How we got here

The constraint on AI infrastructure has been migrating up the stack since 2020, in three distinct eras.

Era 1 (2020-2022): mature-node foundry capacity. The semiconductor lines that produce chips for cars, televisions, and industrial equipment ran short as AI, automotive electrification, and consumer electronics demand converged simultaneously.

Era 2 (2022-2024): advanced packaging and memory. ChatGPT's launch in November 2022 changed the character of demand entirely. Hyperscalers raced to build GPU clusters at scales nobody had planned for, running straight into two new bottlenecks. CoWoS, TSMC's packaging process for stacking high-bandwidth memory alongside a GPU die, and HBM itself, the specialized memory that modern AI chips require. All three HBM suppliers sold out through 2026. CoWoS lagged demand by roughly 20%.

Era 3 (2024-present): power and grid. GPU clusters hit physical power ceilings. The U.S. grid interconnection process, built for industrial customers wanting 50 megawatts, was suddenly receiving 500-megawatt requests from single campuses. FERC rejected Amazon's behind-the-meter nuclear scheme in November 2024. BIS added HBM to export controls a month later. By January 2026, all three constraint layers were binding simultaneously.

2020Era 1,mature-nodefoundry capacityruns short2022BIS restrictsadvanced-nodechip exports,concentratingsupplyEra 2, CoWoSpackaging andHBM becomebindingconstraints2024Era 3, power andgrid emerge asco-constraintMicrosoft restartsThree Mile Islandfor data centerpowerAmazon-TalenFERC rejectionblocksbehind-the-meternuclearBIS adds HBMdirectly to exportcontrols2025Era 4, CoWoS andHBM absorb over90% of advancedpackagingcapacityStargate $500Bformed (capital isnot the constraint)2026Era 5, HBM,CoWoS, power,and grid allbindingsimultaneouslyMeta announces6.6 GW nuclearportfolio (earliestdelivery 2030s)

Figure 1: The constraint migration, 2020-2026.

The underlying mechanism

The cascade does not converge. That is a prediction, not a hedge, and it is rooted in a specific dynamic.

Each time one constraint layer resolves, it does not relieve the system. It accelerates arrivals at the next layer. TSMC's CoWoS capacity is expanding at 80-plus percent compounded annual growth. The effect is to free GPUs for shipment. Those GPUs arrive at data centers that need grid connections. CoWoS relief directly accelerates pressure on the grid queue. Solving the upstream constraint makes the downstream constraint worse, faster.

The grid is the most downstream layer, and it has the slowest service rate in the chain: five to ten years to add meaningful capacity, versus 18 to 30 months for HBM or CoWoS. Product Development Flow (Reinertsen) quantifies what that asymmetry means: at 95% capacity utilization, wait times run 25 times longer than at 75%. The AI chain has been operating near 95% across multiple layers. The grid layer will receive an accelerating wave of arrivals just as its capacity to absorb them is most compressed.

W. Brian Arthur called this mechanism decades before AI was a concern: "Technology A sets up a need for arrangements B; technology C fulfills this, but sets up further needs D and E."

The lens-by-lens read

Theory of Constraints (Goldratt) makes the clearest call. Its cardinal principle: there is always exactly one binding constraint. When HBM and CoWoS resolve, the freed capacity flows into the grid queue. Grid has the longest elevation lead time in the chain (PJM: 7-plus-year average; ERCOT: 233-plus gigawatts queued). It becomes the sole binding constraint, by mechanism. Theory of Constraints (Goldratt) names this System Constraint Recurrence: once you elevate one stage, the next slowest becomes the new bottleneck.

ToC favors Scenario 1 (35%) and argues against Scenarios 2, 3, and 4. Against Scenario 2: multi-constraint deadlock is incoherent in ToC terms, one stage is always slower. Against Scenario 3: demand deceleration does not eliminate the structural grid constraint; grid timelines do not shorten because AI capex pauses. Against Scenario 4: architecture changes address silicon constraints, not grid constraints; the slowest stage persists independently.

Product Development Flow (Reinertsen) agrees on grid dominance but via a stochastic rather than clean sequential path. Its Stochastic Bottleneck describes a constraint that migrates unpredictably as capacity relief arrives unevenly. When HBM eases, arrivals accelerate simultaneously at CoWoS and grid, likely producing a brief CoWoS queue spike in 2026-2027 even as HBM loosens. The August 2025 CoWoS dip to 60% utilization followed by rapid rebound to full-booking within weeks was a live calibration event: the system is in a high-sensitivity stochastic regime.

PDF favors Scenarios 1 and 3, dissents on Scenarios 2 and 4. It favors Scenario 3 because demand deceleration is the one mechanism that returns the system to the stable side of the throughput curve without requiring supply-side grid relief. It dissents on Scenario 4 because architectural transitions create their own new linked queues and no architecture change crosses the exponential-queue inflection point by 2028 at current queue depths.

Complexity Economics (Arthur) agrees on grid dominance and adds that the cascade will not converge by 2028. The economy remakes itself around AI's physical requirements layer by layer; solving grid interconnection reveals transformer backlog, water rights, long-haul transmission, and cooling sub-constraints in sequence. CE favors Scenarios 1 and 4. It favors Scenario 4 because increasing-returns lock-in can be broken by exogenous shocks, and the 2024-2025 window was the critical pre-lock-in period. That window is now largely past, but historical contingency remains live. CE's primary falsifier for Scenario 3: AI demand plateau requires technology paradigm failure, not merely efficiency gains.

Wardley Mapping (Wardley) is the dissenting voice on Scenario 1's severity. Its evolution axis predicts that commodity-adjacent components under constraint pressure generate market workarounds. The grid sits at evolution 0.52: institutionally frozen but technically unremarkable. Hyperscalers have already demonstrated the workarounds: Microsoft's Three Mile Island restart, Meta's nuclear portfolio, brownfield co-location, SMR contracts. Wardley argues these mechanisms will partially circumvent the grid constraint before it reaches maximum binding severity. Against Scenario 2: when HBM is tight, demand routes toward disaggregated memory alternatives; when grid is tight, compute migrates to powered locations. True deadlock requires all substitution pathways to close simultaneously. Wardley favors Scenarios 3 and 4, dissents on Scenarios 1 and 2.

The predictions in depth

Scenario 1: Grid Claims the Cascade (35%)

HBM normalizes by end-2026 as the SK Hynix, Micron, and Samsung investment wave matures. CoWoS closes from roughly 20% to single digits by mid-2027, as TSMC tracks its expansion and ASE and Amkor ($2.5-3 billion in 2026 capex) add overflow capacity. The freed silicon deployment drives accelerating GPU cluster build-outs arriving en masse at the grid queue. By 2027-28, AI datacenter deployment is physically limited by grid-connected megawatts, not chips. Capital accumulates as work in progress in front of the grid bottleneck.

Wardley Mapping dissents on severity: private workarounds will partially route compute around the grid queue, reducing the constraint's binding force below the clean single-bottleneck picture that Theory of Constraints predicts.

Scenario 2: Multi-Front Deadlock (25%)

CoWoS resolves broadly on schedule. But the HBM3E-to-HBM4 transition creates persistent friction: SK Hynix deliberately slows HBM4 ramp to protect HBM3E revenue; Samsung faces quality issues at advanced nodes; BIS controls concentrate global HBM supply onto a smaller Allied-nation base. Memory remains partially binding for frontier training clusters through 2027. Grid independently binds for large-scale inference. No single constraint dominates: HBM limits high-end training, grid limits inference farms.

Theory of Constraints and Wardley Mapping both dissent. ToC: there is always exactly one slowest stage. Wardley: true deadlock requires all substitution pathways to close simultaneously, which it considers unlikely.

Scenario 3: Efficiency Dampens the Cascade (20%)

Foundation model scaling hits diminishing returns before 2027, from data exhaustion, compute-efficient inference architectures, or enterprise adoption saturation. Hyperscaler capex guidance cuts cascade into relaxed demand at every constraint layer. HBM long-term agreements become oversupplied; CoWoS utilization drops below 80%; grid backlog growth slows as new datacenter filings decline. Constraints do not disappear but cease to be investment-limiting.

Complexity Economics and Theory of Constraints both dissent. CE: demand plateau requires technology paradigm failure, not merely efficiency gains. ToC: demand deceleration does not eliminate the structural supply constraint.

Scenario 4: Contingency Reshuffles the Stack (20%)

A high-impact contingency breaks the constraint hierarchy before 2028. Taiwan geopolitical disruption forces rapid diversification away from TSMC toward alternative packaging (FOWLP at Samsung, Intel Foundry), scrambling the constraint map. Or a novel low-power inference architecture achieves threshold adoption and materially reduces energy-per-compute-unit. Or BIS dramatically escalates HBM controls, capping global HBM supply and forcing architectural alternatives. In all variants, increasing-returns lock-in is broken by exogenous shock rather than cascade progression.

Product Development Flow and Theory of Constraints both dissent. PDF: at current queue depths, no architecture change crosses the exponential-queue inflection point by 2028. ToC: architecture changes do not address grid interconnection timelines; the slowest stage persists independently.

The milestones and what they tell you

Milestone Due First check Branch signal
HBM lead times below 26 weeks Q1 2027 Oct 2026 Memory resolving; favors S1 or S3
CoWoS supply gap under 5% Q2 2027 Oct 2026 Packaging clear, grid accelerating; S1 or S2
AI efficiency: >50% GPU-hours saved per model quality unit End 2027 Oct 2026 Demand-side relief gaining traction; S3
Major hyperscaler AI capex cut over 20% End 2027 Oct 2026 Demand deceleration underway; S3
BIS escalates HBM controls to Allied customer impact End 2027 Oct 2026 Multi-front deadlock or S4 pivot more likely
FERC/ISO median datacenter approval below 3 years Mid 2028 Jan 2027 Grid constraint partially relieved; S3 or partial S1 resolution
Power transformer lead times below 80 weeks End 2028 Jan 2027 Sub-constraint inside grid layer resolving

How should each milestone move the weights? If HBM lead times fall below 26 weeks on schedule (Oct 2026 check), S1 and S3 both gain at the expense of S2. If CoWoS clears below 5% while HBM also resolves, S1 gains significantly, grid becomes unambiguously the dominant remaining constraint. If a major hyperscaler cuts AI capex by 20-plus percent, S3 gains substantially at the expense of S1 and S2. If BIS escalates HBM controls to hit Allied customers, S2 and S4 both gain. The grid milestones in 2027-2028 primarily distinguish S1 from partial S1-resolution paths.

2020Era 1, Foundry2022BIS controls / Era2, CoWoS andHBM2024Era 3, power andgrid / MicrosoftTMI / FERC ruling /BIS HBM2025Era 4, over 90%capacityconsumed /Stargate2026Era 5, allconstraintssimultaneous /Meta nuclearOct 2026, firstmilestone check(HBM, CoWoS,capex)2027Q1, HBM leadtime thresholdQ2, CoWoS gapthresholdJan 2027, secondmilestone check(FERC,transformers)End 2027,efficiency, capexcut, BIS escalationtriggers2028Mid, FERC/ISOreform thresholdEnd, transformerlead timethreshold /Forecast closes

Figure 2: History extended with forward checkpoints.

The map

The Wardley map plots each component of the AI compute value chain on two axes. Vertical: visibility to the end user (top = user-facing; bottom = deep infrastructure). Horizontal: evolution (left = bespoke and one-of-a-kind; right = commodity utility).

GenesisCustom-builtProduct (+rental)Commodity (+utility)Visibility (user-facing at top)AI Cloud Services · value-captureGPU / Accelerator Design · value-captureHigh Bandwidth Memory (HBM) · constraintCoWoS Advanced Packaging · constraintLeading-Edge Foundry · gateGrid Interconnection and Power Delivery · constraintEUV Lithography Equipment · gatePower Transformers · constraintconstraintvalue-capturegate

The striking spatial pattern: AI Cloud Services sits near the top right (visibility 0.95, evolution 0.65), and that is where value is captured. HBM and CoWoS sit at the bottom left (visibility 0.45 and 0.40, evolution 0.22 and 0.18), and that is where the constraint has lived through 2025-2026. Grid Interconnection sits unexpectedly far right (evolution 0.52): a mature, commodity-adjacent utility service that is not technically novel but is institutionally frozen. Power Transformers (evolution 0.45) should be unremarkable; 160-week lead times have turned them into a sub-constraint within the grid layer.

GPU and Accelerator Design (visibility 0.70, evolution 0.45) is the value-capture anchor: not the constraint, but where NVIDIA's CUDA ecosystem produces extraordinary margins. Leading-Edge Foundry (evolution 0.30) and EUV Lithography Equipment (evolution 0.35) are gates, not current binding constraints, but where geopolitical tail risk concentrates.

Component Scenario 1 (35%) Scenario 2 (25%) Scenario 3 (20%) Scenario 4 (20%)
GPU / Accelerator Design n/a n/a commoditizes (to 0.58) reprices
High Bandwidth Memory (HBM) commoditizes (to 0.35) locks at current commoditizes (to 0.42) shifts back (to 0.14)
CoWoS Advanced Packaging commoditizes (to 0.30) shifts slightly (to 0.25) commoditizes (to 0.38) disrupted (to 0.12)
Leading-Edge Foundry n/a n/a n/a reprices sharply
Grid Interconnection stays frozen stays frozen shifts toward commodity (to 0.58) n/a
Power Transformers reprices upward n/a n/a n/a

Grid Interconnection stays frozen in the two most probable scenarios (S1 and S2, combined 60%). In a majority of futures, the grid does not move. Everything else piles up in front of it.

The scorecard

No calls on this topic have resolved yet. The track record is empty. This is the panel's opening position.

Framework Stances Favors Against Live probability weight favored
Complexity Economics (Arthur) 4 S1, S4 S3 (primary) 80%
Product Development Flow (Reinertsen) 4 S1, S3 S2, S4 80%
Wardley Mapping (Wardley) 4 S3, S4 S1, S2 40%
Theory of Constraints (Goldratt) 4 S1 S2, S3, S4 35%

Wardley Mapping and Theory of Constraints are in the most exposed positions: Wardley favors only 40% of probability weight, ToC only 35%. Both will receive their first calibration signal at the October 2026 milestone check.

What this means for strategy

If you build data centers: energization timelines belong in the same spreadsheet as fiber, land, and labor. A project that can connect to the grid tomorrow is worth more than one that finishes construction in 2026 and waits until 2031 for power. The 60% combined probability on grid-dominant scenarios says this trade-off is not edge-case analysis. Brownfield industrial sites with existing substations are appreciating in strategic value independent of their real estate characteristics.

If you allocate capital into AI infrastructure: the chip companies are not where scarcity concentrates by 2027-28 in the most probable futures. The constraint is the wire and the transformer at the end of it. Power Transformers have a 160-week backlog and ABB, GE Vernova, Hitachi Energy, and Eaton are not building new plants fast enough. The companies that solve energization have pricing power that silicon companies can only approach when they control private power generation.

If you are a hyperscaler: the Microsoft and Meta private power deals are not energy policy exercises. They are an attempt to route around the dominant constraint before it fully binds. Whether those routes stay open or get closed by FERC and NRC is itself a milestone worth tracking. $500 billion committed to Stargate will produce expensive servers waiting for grid connections if the energization problem does not move.

If you are betting a few years on this analysis: the scenario most likely to make it wrong in a useful way is Scenario 3. AI efficiency breakthroughs that meaningfully reduce compute per quality unit could deflate demand pressure before the grid constraint reaches maximum binding severity. If you see sustained hyperscaler capex reductions at the October 2026 check, the weights shift materially and the analysis requires revision.

The queue is twice the grid. That sentence is the forecast.