AI Infra
0%
Part IX · Chapter 68

Powering It: Time-to-Power as the Binding Constraint

AuthorChangkun Ou
Reading time~12 min

Chapter 62 ended at the rack. The next constraint appears where the rack plugs in. In 2026 the binding constraint on AI capability moved below the chip to the physical layer of electricity. The scarce quantity there is not generation capacity but time-to-power: a frontier campus now self-generates rather than waits for the grid, a rack drawing as much power as dozens of homes forces liquid cooling as the default, and the data center has become something the grid operator must actively manage rather than merely serve.

2026-06-21T21:25:16.083507 image/svg+xml Matplotlib v3.11.0, https://matplotlib.org/
Figure 68.1. Schematic of the time-to-power gap. AI load demand can grow faster than delivered grid capacity, so power arrives on a slower clock than chips. Idealized curves, not measured data.

The previous two chapters followed a constraint arrow down: from the bytes a chip can move, to the interposer and memory stacks that gate how many chips exist. The path now reaches the bottom of the physical stack. Below the chip and below the package is the electricity, and electricity binds on a timescale longer than either.

The chips arrive in months, the megawatts in years

A frontier training campus now needs on the order of a gigawatt of dedicated power, comparable to a mid-sized city or a nuclear unit. The headline fear is that the world cannot generate enough electricity, but that is not the 2026 constraint. The constraint is time. A chip order is filled in months; the power behind it arrives in years, because three lead-time queues sit in series, and a campus is energized only when the slowest clears.

The first queue is grid interconnection. Connecting a large new load or generator to the transmission system requires a study-and-approval process that, in the United States, backs up for years: the active interconnection queue holds on the order of two thousand gigawatts of projects, far more than the country's entire installed capacity, with typical waits measured in multiple years (Rand et al. 2025). The pressure is visible in a single market: in Texas, large-load interconnection requests jumped from roughly 63 GW at the end of 2024 to roughly 226 GW a year later, the large majority of it data centers. Behind it sits long-lead equipment. A large power transformer, the device that steps grid voltage down for a campus, runs lead times around 128 weeks against a structural supply deficit (Wood Mackenzie 2025). Last come the prime movers: the gas turbines that on-site generation depends on are themselves backed up, with a single major manufacturer's backlog passing 80 GW in 2025 (Cardwell 2025) and, counting reservations, reaching about 100 GW by early 2026, effectively sold out through 2030 (Martucci 2026).

chip accelerators ~months live campus energized chip->live arrives early, waits q1 grid interconnection multi-year queue q1->live q2 long-lead equipment transformers ~128 weeks q2->live q3 prime movers turbines sold out into 2030 q3->live
Figure 68.2. Time-to-power as three serialized queues. A frontier campus is energized only when the slowest clears, so the binding constraint is calendar time, not generation capacity. The chip, by contrast, ships in months.
Constraint arrow

The serialized lead times of interconnection, transformers, and turbines cap how fast cluster scale can grow, independent of how much capital or how many chips a buyer has. This is the arrow that most surprises a reader coming up from the silicon chapters: more money and more accelerators do not buy a faster build once the binding queue is a transformer factory or a turbine backlog. The megawatt, not the GPU, now sets the schedule (cross-ref Chapter 76).

Going behind the meter

Because the interconnection queue is the longest pole, the largest builders stopped waiting in it. The defining operational move of 2026 is self-generation behind the meter: standing up power on the campus itself rather than drawing it through a grid connection that is years away.

On-site gas is the fast option. Operators have built campuses around dozens of gas turbines fired on site, with the largest planned builds reaching one to several gigawatts. It is the quickest path to bulk power, and it collides directly with the turbine backlog and with air-quality law, which makes it the chapter's sharpest contested point. For firm power, the answer is nuclear. The most visible deal of the cycle restarts a shuttered reactor under a long-term power-purchase agreement to serve a hyperscaler, targeting a return to service in 2027 and backed by a federal loan (NucNet 2025); other restarts and uprates add firm gigawatts on a similar horizon. The regulatory path is not clean: a marquee co-location deal that would have placed a data center directly behind a nuclear plant's meter was rejected by federal regulators and had to be restructured as a conventional grid-connected agreement, a reminder that behind-the-meter is a legal posture as much as an engineering one.

small modular reactor (SMR) come up constantly as the third option, and what matters is placing them correctly on the timeline. The contracted pipeline grew large across 2025 and 2026, with multiple multi-gigawatt agreements signed, but no commercial unit dedicated to AI is operating, and the first are not expected before roughly 2030. They are a real hedge for the next decade, not a source that relieves the 2026 to 2028 crunch. Confusing the contract with the kilowatt is the most common error in this part of the discourse.

A quieter inversion runs underneath all of these. Each option above brings power to the compute: fix the campus, then generate or contract the megawatts to energize it. The opposite move is to bring the compute to the power, siting the machines where energy is already stranded. Crusoe built its first business this way, dropping modular data centers at oil wells to consume gas that would otherwise be flared, turning a waste stream into compute at the wellhead, and has since grown into an energy-first, vertically integrated builder of gigawatt-scale AI campuses on stranded and renewable power (Yee and Lochmiller 2025). The founding logic still marks the cheapest power on the map, stranded, curtailed, or remote energy that no one else is positioned to use, and it turns the build decision into a question of where the spare electricity already sits.

The rack becomes a heat problem

The power problem reappears as a heat problem inside the rack. A current flagship rack draws roughly 120 kW, with observed figures past 130 kW, which is several times what forced-air cooling can remove from a rack even with aggressive engineering. Air, the default for the entire prior history of the data center, simply cannot carry the heat away. Direct-to-chip liquid cooling, where coolant runs through cold plates pressed onto the accelerators, became the non-optional default for this class of hardware, and immersion, submerging the boards in dielectric fluid, is the next tier for the densities coming after. The facility-efficiency frontier, meanwhile, has flattened: the industry-average power usage effectiveness (PUE), the ratio of total facility power to IT equipment power, has sat near 1.54 for years while the best hyperscalers run near 1.1, which means the easy facility-level gains are spent and the remaining efficiency must come from the silicon, through tokens per watt (the output tokens produced per watt of power), and from the grid.

Constraint arrow

Rack power density past the air-cooling ceiling dictates the thermal architecture of the whole facility: a 120 kW rack forces direct-to-chip liquid as the default, which reshapes the data hall, the plumbing, and the supply chain for cooling gear. This is the per-rack power of Chapter 62 reaching down into the building (cross-ref Chapter 82).

The data center the grid must manage

The newest frontier is the one the earlier framing missed entirely: at gigawatt scale, an AI data center is not a docile load but a hazard to the stability of the grid it sits on. A large training cluster can shed or pick up a thousand megawatts in seconds when a job stops or a fault trips protection, and a swing that large, that fast, is something the bulk power system was not designed to absorb. After a series of such events, including a load drop above a gigawatt during a fault on a major interconnection, the North American reliability authority issued a rare high-severity alert in May 2026, and a wave of research now studies the wide-area oscillations and fault ride-through behavior that large AI loads induce. The regulatory layer is moving in parallel: at the end of 2025 federal regulators directed the largest grid operator to write rules for data-center-at-power-plant co-location, with compliance beginning in 2026, which is the legal scaffolding underneath the behind-the-meter moves above.

The same scale creates a cost fight. Someone must pay for the transmission upgrades and the capacity a gigawatt campus demands, and as data-center load drove capacity-auction prices to records, the question of whether ordinary ratepayers or the data centers themselves bear the cost became a live battle at state regulators. A more cooperative path is also maturing: treating the cluster as a flexible, curtailable load that can ramp down briefly to relieve the grid lets far more compute interconnect without new firm generation, an idea that moved from study to pilot during this period. Whether AI load is a threat to be managed or a flexible resource to be harnessed is not yet settled.

grid dc gigawatt AI campus hazard hazard: GW load swings in seconds NERC alert, oscillation risk dc->hazard flex resource: curtailable load interconnect without new firm power dc->flex cost cost fight: who pays for upgrades (ratepayers vs operators) dc->cost
Figure 68.3. The gigawatt campus as a two-sided object: a reliability hazard the grid must absorb (fast multi-gigawatt load swings) and, potentially, a flexible resource that curtails to interconnect faster. Which face dominates is the chapter's open question.

How much, and measured how

The aggregate numbers frame the stakes. Global data-center electricity ran near 415 TWh in 2024, about 1.5 percent of world demand, and is projected to roughly double toward 945 TWh by 2030, with AI-optimized load growing fastest (International Energy Agency 2025). In the United States the most-cited study put data centers at 176 TWh in 2023, rising to a range of 325 to 580 TWh, between roughly 7 and 12 percent of national electricity, by 2028 (Shehabi et al. 2024). The width of that range is the real point: the future depends on efficiency gains and build-rates nobody can yet pin down.

Per query, the picture is contested precisely because the boundary is contested. A vendor's full-stack accounting put the median text prompt near 0.24 Wh of energy, including idle capacity and overhead, and claimed a large drop over the preceding year (Elsworth et al. 2025). Critics note that favorable boundary choices, and the exclusion of training amortization, make cross-model comparisons unreliable, and that no audited standard exists. The energy-per-token number is real and useful and not yet trustworthy as a comparison.

What's contested
  • Does power or silicon bind first? A power-limited camp argues chips are available and megawatts are not, so tokens per watt sets revenue. A supply-chain camp argues the binding link migrates year to year between HBM, packaging, transformers, and turbines, and in 2026 it is electrical equipment rather than chips. They agree it is no longer FLOPs.
  • Are demand forecasts overblown? Skeptics hold that load forecasts are chronically overstated and that efficiency absorbs much of the growth (Koomey et al. 2025). Utilities and several research groups hold that the build-out is real and large. The truth turns on efficiency and build-rate assumptions that are themselves the disagreement.
  • Is nuclear and SMR power real or a hedge? Restarts and uprates are contracted firm gigawatts arriving 2027 to 2028. New small reactors are a ~2030 proposition at the earliest, so they cannot relieve the near-term crunch however large the signed pipeline grows.
  • Is on-site gas a bridge or a liability? One camp calls behind-the-meter gas the only way to gigawatts before 2030. The other points to air-quality litigation and broken climate pledges. The fight escalated rather than resolved, drawing in federal intervention on national-security grounds.
  • Narrow or full-stack energy accounting? Without an audited standard, every published per-query figure is a boundary choice, and the boundary is the argument.

When megawatts decide models

Here capability, efficiency, and trust all land on the calendar. A model that could be trained if the power existed is not trained until the slowest queue clears, which makes the frontier partly an infrastructure-delivery frontier and puts capability itself on a delivery schedule. Efficiency, meanwhile, stopped being a facility metric and became a silicon and grid metric, because the data hall's easy gains are spent and the remaining ones live in tokens per watt and in flexible load. Trust, at this layer, is literal: the grid's trust in a load that can swing a gigawatt in seconds, and a community's trust in a campus burning gas next door. The constraint arrow that began with bytes inside a chip now reaches the electrical grid and the permitting office, and it points back up the whole stack: the megawatt dictates the model. The next chapter goes back inside the cluster the power feeds, and asks why a machine of a hundred thousand accelerators is failing somewhere almost all the time.

Further reading

  • International Energy Agency, “Energy and AI” (the global baseline: ~415 TWh in 2024 toward ~945 TWh by 2030), 2025. iea.org
  • Shehabi et al., “2024 United States Data Center Energy Usage Report” (the canonical US load study, 176 TWh in 2023 to 325--580 TWh by 2028), 2024. escholarship.org
  • Rand et al., “Queued Up: 2025 Edition. Characteristics of Power Plants Seeking Transmission Interconnection” (the interconnection-queue data behind the time-to-power thesis), 2025. emp.lbl.gov
  • Elsworth et al., “Measuring the Environmental Impact of Delivering AI at Google Scale” (the full-stack 0.24 Wh-per-prompt figure and the measurement-boundary debate), 2025. arXiv:2508.15734
    Google measures the full-stack energy, carbon, and water footprint of Gemini Apps inference in production, finding the median text prompt consumes 0.24 Wh and showing a 44x emissions reduction over one year.
  • Cardwell, “GE Vernova Expects to End 2025 with an 80-GW Gas Turbine Backlog That Stretches into 2029” (the prime-mover and long-lead-equipment queues), 2025. utilitydive.com
    GE Vernova expects to close 2025 with an 80-GW gas turbine backlog stretching to 2029, with reservations projected to sell out through 2030 by end of 2026.
  • Wood Mackenzie, “Power Transformers and Distribution Transformers Will Face Supply Deficits of 30% and 10% in 2025” (the prime-mover and long-lead-equipment queues), 2025. woodmac.com
  • Koomey et al., “Electricity Demand Growth and Data Centers: A Guide for the Perplexed” (the demand-skeptic position), 2025. bipartisanpolicy.org
  • Yee & Lochmiller, “How Crusoe Powers and Transforms AI with Stranded Energy” (siting compute at stranded energy, the inverse of bringing power to the campus), 2025. mckinsey.com
    In a McKinsey interview, Crusoe CEO Chase Lochmiller describes the company's energy-first model: it began by placing modular data centers at oil wells to consume otherwise-flared gas, then grew into a vertically integrated builder of large AI campuses sited on stranded and renewable power.
  • Lochmiller & Agrawal, “Economics of the AI Supercycle, Class 3” (Crusoe's CEO on financing, siting, and powering AI data centers), 2026. youtube.com
    Class 3 of Stanford's MS&E 435 seminar on AI-stack economics: a conversation with Crusoe co-founder and CEO Chase Lochmiller on financing, siting, and powering large AI data centers from an energy-first vantage point.
  • Martucci, “GE Vernova Gas Turbine Backlog Hits 100 GW as Prices Rise” (the backlog-plus-reservations figure a year on), 2026. utilitydive.com
    GE Vernova's gas turbine backlog plus reservations reached 100 GW in Q1 2026, with the company expecting reservations sold out through 2030 by year-end.

Comments

Log in to comment