In September 2024, Microsoft signed a deal to restart Three Mile Island, a nuclear plant that had been shuttered since 2019. The press framed it as energy policy. It was actually an infrastructure problem: Microsoft needed power for data centers and the grid could not deliver it on any timeline that worked.
Four months later, Meta announced it was pursuing 6.6 gigawatts of nuclear generation across multiple sites, with the earliest delivery after 2030. The AI build-out those plants are meant to support is happening now.
That gap is the constraint. Not the chips, not the money. The wire from the substation to the server room.
The panel's view
| Scenario | Probability |
|---|---|
| Grid Claims the Cascade: power/interconnection dominant by 2028 | 35% |
| Multi-Front Deadlock: HBM friction + grid bind simultaneously | 25% |
| Efficiency Dampens the Cascade: AI demand decelerates before 2028 | 20% |
| Contingency Reshuffles the Stack: architecture or geopolitical pivot | 20% |
No calls have resolved. This is the panel's opening position.
Three of four analytical frameworks converge on grid as the dominant binding constraint by 2027-28. The fourth agrees on direction but dissents on severity. The 35% plus 25% combined means a 60% probability that power is part of the constraint picture by 2027-28 regardless of how silicon resolves.
How we got here
The constraint on AI infrastructure has been migrating up the stack since 2020, one layer at a time.
It started with mature-node foundry capacity. In 2020 and 2021, as AI, automotive, and consumer electronics demand hit simultaneously, the foundry lines that produce chips for everything from cars to TVs ran short. That was the first era of the cascade.
Then ChatGPT shipped in November 2022 and everything changed scale. Within months, hyperscalers were racing to build GPU clusters nobody had planned for. That race ran straight into two new bottlenecks: CoWoS packaging (TSMC's process for stacking high-bandwidth memory chips directly alongside a GPU die) and HBM itself. All three HBM suppliers sold out through 2026. CoWoS supply lagged demand by roughly 20%.
By 2024, a third layer emerged. GPU clusters were large enough and numerous enough that their power demand exceeded what existing grid infrastructure had anticipated. The U.S. grid interconnection process, designed for industrial customers wanting 50 megawatts, was suddenly receiving requests for 500 megawatts from a single campus. In November 2024, FERC, the Federal Energy Regulatory Commission, rejected Amazon's plan to connect a data center directly behind a nuclear plant's meter. The following month, BIS added HBM to export controls. By January 2026, all three layers were binding simultaneously.
Figure 1: The constraint migration, 2020-2026.
The underlying mechanism
Here is what is non-obvious. Each time one layer of the constraint resolves, it does not relieve the system. It accelerates the next constraint.
When TSMC expands CoWoS capacity at 80-plus percent compounded annual growth, the effect is to free GPUs for shipment. Those GPUs arrive at data centers that need grid connections. The packaging relief directly accelerates pressure on the grid queue. Solving the upstream constraint makes the downstream constraint worse, faster. W. Brian Arthur called this mechanism decades before AI was a concern: "Technology A sets up a need for arrangements B."
Product Development Flow (Reinertsen) quantifies the dynamic. At 95% capacity utilization, average wait times run 25 times longer than at 75%. The AI chain has been operating near 95% across multiple layers. The grid, the most downstream layer, has the slowest service rate: five to ten years to add meaningful capacity versus 18 to 30 months for HBM or CoWoS. That asymmetry in service rates is what makes the grid dominant. When an upstream queue clears, the grid receives an accelerating wave of arrivals precisely when its capacity to absorb them is most compressed.
How each framework reads the situation
Theory of Constraints (Goldratt) makes the clearest call. Its cardinal principle: there is always exactly one binding constraint. The system's output is governed entirely by the slowest stage. When HBM and CoWoS resolve, the freed capacity flows into the grid queue. Grid has the longest elevation lead time in the chain (PJM: 7-plus-year average, ERCOT: 233-plus gigawatts queued). It becomes the sole binding constraint, by mechanism.
Theory of Constraints (Goldratt) names this System Constraint Recurrence: once you elevate one stage, the next slowest becomes the new bottleneck. There is also a sub-constraint inside the grid layer. Power transformers, the large units above 100 megavolt-ampere at the grid-to-datacenter interface, stretched from 50-week lead times in 2021 to 160-plus weeks by 2026. Even projects that win interconnection approval still face a multi-year equipment wait. ToC calls this a Policy Constraint: an institutional arrangement that persists and resists alteration even when the physical need is clear.
ToC favors Scenario 1 (35%) and argues against the other three. Multi-front deadlock (Scenario 2) is incoherent in ToC's framework: one stage is always slower by at least a marginal degree. Demand deceleration (Scenario 3) does not eliminate the structural constraint; grid timelines do not shorten because AI capex pauses. Architectural disruption (Scenario 4) addresses silicon constraints, not grid constraints; the slowest stage persists independently.
Product Development Flow (Reinertsen) agrees on grid dominance but via a stochastic rather than clean sequential path. Its Stochastic Bottleneck describes a constraint that migrates unpredictably as capacity relief arrives unevenly and demand arrives in bursts. When HBM eases, arrivals accelerate simultaneously at CoWoS and grid, likely producing a brief CoWoS queue spike in 2026-2027 even as HBM loosens. The August 2025 CoWoS dip to 60% utilization followed by rapid rebound to full-booking within weeks was a live calibration event: the system is in a high-sensitivity stochastic regime.
PDF favors Scenarios 1 and 3, dissents on Scenarios 2 and 4. It favors Scenario 3 because demand deceleration is the one mechanism that returns the system to the stable side of the throughput curve without requiring supply-side grid relief. It dissents on Scenario 4 because architectural transitions create their own new linked queues and at current queue depths no architecture change crosses the exponential-queue inflection point by 2028.
Complexity Economics (Arthur) agrees on grid dominance and adds that the cascade will not converge by 2028. The economy remakes itself around AI's physical requirements layer by layer; solving grid interconnection reveals transformer backlog, water rights, long-haul transmission, and cooling sub-constraints in sequence. CE favors Scenarios 1 and 4. It favors Scenario 4 because increasing-returns lock-in can be broken by exogenous shocks, and the 2024-2025 window was the critical pre-lock-in period. CE's primary falsifier for Scenario 3: AI demand plateau requires technology paradigm failure, not merely efficiency gains. The increasing-returns dynamics of CUDA, cloud APIs, and ML talent create positive feedbacks that make demand deceleration unlikely without a full paradigm shift.
Wardley Mapping (Wardley) is the dissenting voice on Scenario 1's severity. Its evolution axis predicts that commodity-adjacent components under constraint pressure generate market workarounds. The grid sits at evolution 0.52: institutionally frozen but technically unremarkable. Hyperscalers have already demonstrated the workarounds: Microsoft's Three Mile Island restart, Meta's nuclear portfolio, brownfield co-location, SMR contracts. Wardley argues these mechanisms will partially circumvent the grid constraint before it reaches maximum binding severity, reducing its force below the clean single-bottleneck picture ToC predicts. Against Scenario 2, Wardley expects that when HBM is tight, demand routes toward disaggregated memory alternatives; when grid is tight, compute migrates to powered locations. True deadlock requires all substitution pathways to close simultaneously. Wardley favors Scenarios 3 and 4, dissenting on Scenarios 1 and 2.
Methodological tension: Theory of Constraints (Goldratt) predicts a clean single-constraint migration; Product Development Flow (Reinertsen) predicts a messy stochastic bounce between nodes. These two views generate different predictions about the CoWoS-to-grid transition in 2026-2027. CoWoS utilization data through 2026 will help distinguish which mechanism is operating.
The milestones
| Milestone | Due | First check | Branch signal |
|---|---|---|---|
| HBM lead times below 26 weeks | Q1 2027 | Oct 2026 | Memory resolving; favors S1 or S3 |
| CoWoS supply gap under 5% | Q2 2027 | Oct 2026 | Packaging clear, grid accelerating; S1 or S2 |
| AI efficiency: >50% GPU-hours saved per model quality unit | End 2027 | Oct 2026 | Demand-side relief gaining traction; S3 |
| Major hyperscaler AI capex cut over 20% | End 2027 | Oct 2026 | Demand deceleration underway; S3 |
| BIS escalates HBM controls to Allied customer impact | End 2027 | Oct 2026 | Multi-front deadlock or S4 pivot more likely |
| FERC/ISO median datacenter approval below 3 years | Mid 2028 | Jan 2027 | Grid constraint partially relieved; S3 or partial S1 resolution |
| Power transformer lead times below 80 weeks | End 2028 | Jan 2027 | Sub-constraint inside grid layer resolving |
Figure 2: History extended with forward checkpoints.
The map
The Wardley map plots each component of the AI compute value chain on two axes. Vertical: visibility to the end user (top = user-facing; bottom = deep infrastructure). Horizontal: evolution (left = bespoke and one-of-a-kind; right = commodity utility).
The striking spatial pattern: AI Cloud Services sits near the top right (visibility 0.95, evolution 0.65), and that is where value is captured. HBM and CoWoS sit at the bottom left (visibility 0.45 and 0.40, evolution 0.22 and 0.18), and that is where the constraint has lived through 2025-2026. Grid Interconnection sits unexpectedly far right (evolution 0.52): a mature, commodity-adjacent utility service that is not technically novel but is institutionally frozen. Power Transformers (evolution 0.45) should be unremarkable; 160-week lead times have turned them into a sub-constraint within the grid layer.
GPU and Accelerator Design (visibility 0.70, evolution 0.45) is the value-capture anchor: not the constraint, but where NVIDIA's CUDA ecosystem produces extraordinary margins. Leading-Edge Foundry (evolution 0.30) and EUV Lithography Equipment (evolution 0.35) are gates, not current binding constraints, but where geopolitical tail risk concentrates.
| Component | Scenario 1 (35%) | Scenario 2 (25%) | Scenario 3 (20%) | Scenario 4 (20%) |
|---|---|---|---|---|
| GPU / Accelerator Design | n/a | n/a | commoditizes (to 0.58) | reprices |
| High Bandwidth Memory (HBM) | commoditizes (to 0.35) | locks at current | commoditizes (to 0.42) | shifts back (to 0.14) |
| CoWoS Advanced Packaging | commoditizes (to 0.30) | shifts slightly (to 0.25) | commoditizes (to 0.38) | disrupted (to 0.12) |
| Leading-Edge Foundry | n/a | n/a | n/a | reprices sharply |
| Grid Interconnection | stays frozen | stays frozen | shifts toward commodity (to 0.58) | n/a |
| Power Transformers | reprices upward | n/a | n/a | n/a |
Grid Interconnection stays frozen in the two most probable scenarios (S1 and S2, combined 60%). In a majority of futures, the grid does not move. Everything else piles up in front of it.
Confidence and caveats
No calls on this topic have resolved yet. The track record here is empty; this is the panel's opening position, and scenario weights will be updated as milestones trigger.
The highest-confidence call is the directional one: power and grid are part of the constraint picture by 2027-28 (60% combined). The weaker distinction is clean grid dominance (S1, 35%) versus a messier two-front bind (S2, 25%). The CoWoS-to-HBM-to-grid transition path is where the biggest modeling disagreement sits, between Theory of Constraints and Product Development Flow.
Scenario 3 (efficiency dampens demand, 20%) is the primary scenario in which this analysis is wrong in a useful way. AI efficiency breakthroughs that materially reduce compute per quality unit could deflate demand pressure before the grid constraint reaches maximum severity. DeepSeek-R1's compute efficiency results in early 2025 moved this from tail-risk to live hypothesis; the question is whether efficiency gains persist and compound or plateau.
The queue is twice the grid. That sentence is the forecast.