Energy and Power as the Binding Constraint
Why a grid connection measured in megawatts, not a budget measured in dollars, now caps how much AI compute can be built, how PUE converts a power limit into an accelerator count, and why the demand forecasts disagree by a factor of two.
Data centres used about 415 TWh of electricity in 2024, roughly 1.5 percent of global consumption, and the International Energy Agency's base case has that more than doubling to around 945 TWh by 2030, slightly more than Japan uses today (IEA, 2025, Energy and AI, executive summary). The same report estimates that around 20 percent of planned data centre projects could be delayed by grid constraints, and for anyone building capacity that second number matters more. A cluster can be financed in a quarter. The substation and the generation behind it take years.
Capex and depreciation and the cost floor of inference treat compute as a dollar problem. This concept treats it as a watt problem, because in many markets that constraint binds first.
From a megawatt cap to an accelerator count
A site's grid connection fixes the maximum facility power \(P_{\text{facility}}\). Only part of it reaches the IT equipment. Power usage effectiveness is the ratio
where \(E_{\text{facility}}\) is all energy entering the site over a period and \(E_{\text{IT}}\) is the energy delivered to servers, storage and network. Everything above 1 is cooling, power conversion losses and lighting. If \(p\) is the average all-in IT power per accelerator (the chip plus its share of CPUs, memory, switches and NICs), the number of accelerators a connection can carry is
Take a 100 MW connection. The Uptime Institute's 2025 survey put the weighted average annual PUE of respondent facilities at 1.54, essentially flat for six years (Uptime Institute, 2025, Global Data Center Survey). Google reports a fleet-wide trailing twelve-month PUE of 1.09 for 2024 (Google Data Centers, Power usage effectiveness). At 1.54 the IT budget is 64.9 MW; at 1.09 it is 91.7 MW. A liquid-cooled rack of 72 current-generation accelerators is specified at roughly 120 kW, about 1.7 kW per accelerator all-in, so the same grid connection carries roughly 38,000 accelerators at the industry average and roughly 54,000 at hyperscaler efficiency.
That 41 percent difference is the point. Under a power cap, PUE is a capacity metric, not an operating-cost metric, and a cooling design decision changes how much compute a site can ever hold.
The energy bill is modest by comparison. Running 100 MW at an average 80 percent draw for a year consumes \(100 \times 8{,}760 \times 0.8 \approx 700{,}800\) MWh; at an assumed $80 per MWh that is about $56 million, a fraction of the hardware's price. Power binds long before it is a large cost, which is why sites with connections already in place command premiums.
The queue in front of the grid
In the United States, new generation and storage must pass through an interconnection study before connecting. At the end of 2025 about 8,200 active projects, 1,312 GW of generation and 749 GW of storage, were waiting, and for projects that reached commercial operation in 2025 the median time from interconnection request was over five years (Berkeley Lab, 2026, Queued Up: 2026 Edition). Those figures describe supply trying to connect, not data centres, which face their own load studies and transmission upgrades; but the supply queue sets how fast new generation can serve them. Its mix is shifting too: active gas capacity in the queue rose 86 percent in 2025 while solar, wind and storage fell.
The planning consequence is a mismatch of clocks. Accelerator generations turn over in about two years; power infrastructure in five to ten. A campus sized for today's rack density may be obsolete in cooling before it is fully energised.
Where the forecasts disagree
Projections span a wide range, and the spread is itself the finding. Berkeley Lab estimated US data centres used 176 TWh in 2023, 4.4 percent of national electricity, and projected 325 to 580 TWh by 2028, between 6.7 and 12 percent (Berkeley Lab, 2024, United States Data Center Energy Usage Report). The IEA's own scenarios for 2035 run from 700 to 1,700 TWh globally.
Sceptics point to history. Masanet and colleagues found global data centre energy use rose only about 6 percent between 2010 and 2018 while compute instances rose about 550 percent, as virtualisation, hyperscale consolidation and PUE gains absorbed demand, and they challenged analyses that had predicted doubling (Masanet et al., 2020, Recalibrating Global Data Center Energy-Use Estimates, Science 367(6481)). The counter-argument is that the easy wins are spent: PUE has plateaued for the average operator, and accelerated servers draw far more per rack than the CPU fleet that efficiency gains once shrank. Whether efficiency or induced demand wins is not settled by any current evidence.
When it breaks
PUE is silent about the IT side. It measures overhead, not useful work per IT watt: a site at 1.1 running idle accelerators spends more energy per token than a site at 1.4 running them hot. It also ignores water, since evaporative cooling lowers PUE by consuming it.
Nameplate power is not drawn power. Provisioning to thermal design power strands capacity; provisioning to measured averages risks tripping breakers when thousands of devices ramp together in a synchronised training step. Power capping and oversubscription recover capacity at the cost of occasional throttling.
Annual energy is the wrong unit for the grid. Utilities plan for peaks. A training cluster running flat out is a large, inflexible block of peak load, and shifting batch work away from peaks needs tariffs and contracts most AI deployments do not yet have.
Local constraints dominate global averages. A national share of 4 percent can be a double-digit share in one utility territory. The binding constraint is a specific substation, not a national statistic.
7 flashcards for this concept
Click a card to reveal the answer.