The AI Compute Buildout Is a Power Problem Now
Interconnection queues and transformer backlogs set the pace, not accelerator orders. Where the buildout stands in 2026 and what would change its trajectory.
In this story 5 sections
The AI compute buildout is now constrained by electricity and grid connection rather than chip supply. Datacenters used roughly 415 terawatt-hours in 2024, about 1.5 percent of global electricity, and AI-optimized capacity is the fastest-growing share. Power procurement and interconnection queues, not accelerator orders, set how fast new capacity comes online.
Ask an infrastructure team what is holding up their next datacenter and you will rarely hear chips. You will hear substations, transformers, and a utility interconnection queue measured in years.
This piece covers where the AI compute buildout stands in 2026, why power became the binding constraint, what the demand forecasts actually say, and which risks would change the trajectory. It is written for people tracking AI infrastructure as an industry rather than a product.
Where the AI Compute Buildout Stands in 2026
Capital spending on AI infrastructure by the largest cloud providers has risen sharply for three consecutive years, and 2026 commitments are the largest yet announced. The spending is concentrated among a handful of companies.
Combined capital expenditure across the largest hyperscalers has been tracking well past $200 billion annually as of 2026. Three years earlier, it was under $100 billion. That growth rate outpaces the physical infrastructure needed to support it. Dollars can be committed in a quarter. A substation upgrade cannot.
Most of it goes to AI-specific capacity: accelerators, the networking to connect them, and the power and cooling to keep them running. Conventional server capacity is a shrinking share of new build.
Leasing versus building splits the field. Some operators sign long-term capacity agreements with colocation providers, others build owned campuses, and the choice trades speed of deployment against cost per megawatt over a decade.
The strategic logic is two-sided. Training frontier models requires enormous clusters, and serving deployed products requires steady inference capacity that grows with usage. Those are different workloads with different hardware profiles, and inference is the one that scales with revenue.
Serving costs also depend on what the models require, since a sparse model needs enough memory for every parameter even when it computes with few. Our explainer on how sparse architectures split total from active parameters covers why capacity planning tracks the larger number.
Third-party analysis of the buildout's assumptions, including the returns required to justify current spending, is published in the research notes at Goldman Sachs Insights. The open question in nearly all of these analyses is not whether demand exists but whether it arrives on the schedule the capital assumes.
Why Is Power the Bottleneck Instead of Chips?
Power is the bottleneck because electricity generation, transmission, and grid connection all move on multi-year timelines that money cannot compress, while accelerator supply can respond to capital within a year. A large AI datacenter draws power comparable to a small city, and that power has to be delivered before the first rack is energized.
Three separate bottlenecks stack up. Generation capacity has to exist. Transmission has to reach the site. And the interconnection agreement with the utility has to clear a queue that runs years long in many US regions.
Cooling is the quiet third constraint. High-density accelerator racks exceed what air cooling handles well, and liquid cooling changes the building design, the water requirements, and the siting conversation with a local utility.
Equipment lead times compound it. High-voltage transformers and switchgear have multi-year backlogs, and no amount of capital shortens a manufacturing queue that is already full.
We have tracked orders at Emergent Wire where a utility-scale transformer quoted at roughly a year's lead time in 2021 now runs three to four years. The same handful of global manufacturers supply both grid utilities and every hyperscaler at once. Paying a premium moves a project up the queue only slightly. The factories are running near capacity regardless of price.
The scale is documented. Datacenter electricity consumption was around 415 terawatt-hours in 2024, roughly 1.5 percent of global consumption, and is projected to roughly double by 2030, according to the International Energy Agency (IEA) analysis of energy and AI. Electricity demand from AI-optimized datacenters specifically is projected to grow several times faster than the overall figure.
| Constraint | Typical lead time | Can capital compress it? |
|---|---|---|
| Accelerator supply | Months to a year | Yes, partly |
| Building construction | 1 to 2 years | Yes |
| Grid interconnection | Multiple years | Rarely |
| Transformers, switchgear | Multiple years | No |
| New generation capacity | Years | No |
What the Demand Forecasts Actually Say
Forecasts in this area carry wide error bars, and the honest summary is a range rather than a number.
Two things drive the spread. Nobody knows how quickly inference demand grows as products mature, and nobody knows how much efficiency work will land per year. Small differences in either assumption compound into very different 2030 numbers.
Regional concentration is the part that gets understated. National electricity figures look manageable while specific grids absorb a large share of new load, and that is where price effects and connection moratoriums appear first.
Northern Virginia is the clearest US example. Data Center Alley already draws a substantial share of the region's grid capacity. The local utility has flagged multi-year waits for new large-load interconnections in parts of the territory. A developer proposing a new campus there today is often quoted a connection date years out, no matter how quickly the building itself could be finished.
Efficiency runs the other way. Cost per token has fallen steadily across model generations through better hardware, sparse architectures, quantization, and serving optimizations. The mechanism behind one of those is covered in our explainer on how mixture of experts models cut compute per token.
Efficiency does not settle the question, though. Cheaper inference historically increases usage rather than reducing total consumption, and that rebound is why per-token efficiency gains have not slowed aggregate demand growth.
Water gets less attention than power and matters locally. Evaporative cooling consumes real volumes in regions that are already stressed, and several proposed sites have been renegotiated over water rather than electricity.
A large facility using evaporative cooling can consume several million gallons of water a day at peak load. That is comparable to a small town's municipal supply. In drought-prone regions of the American Southwest, that single line item has sunk site selections that otherwise cleared every power and permitting hurdle.
US-specific generation and consumption data, which is the right resolution for the regional question, comes from the U.S. Energy Information Administration (EIA). Its state-level series show the load growth concentrating in a small number of datacenter corridors rather than spreading evenly.
What Would Change the Trajectory
Four developments would meaningfully alter the buildout, listed from most to least likely in our reading at Emergent Wire.
- Demand arrives slower than the capital assumed. Utilization is what turns a datacenter into a return. Capacity built for demand that shows up two years late is expensive idle equipment.
- Efficiency outruns usage growth. A step change in inference cost per useful task would reduce required capacity, though history suggests usage expands to absorb it.
- Local opposition and grid policy tighten. Several jurisdictions have already slowed or conditioned large datacenter connections, and that constraint spreads faster than generation gets built.
- Workloads shift off centralized infrastructure. Routine inference moving to user devices removes some datacenter load, a pattern covered in our piece on where small models genuinely win.
These interact. Slower demand and tighter grid policy arriving together would stall projects already under construction, which is a different and more expensive problem than never having started them.
A half-built campus with signed power contracts but falling utilization is worse for a balance sheet than a canceled project. The fixed costs of land, permits, and partial construction are already sunk. The revenue that was supposed to cover them has not materialized. That is the scenario most worth watching heading into 2027.
Utilization is the number to watch, and it is the one least often disclosed. Announced capacity tells you what was committed. Revenue per unit of deployed compute tells you whether it was needed.
Emergent Wire looks for this in quarterly disclosures whenever a hyperscaler breaks out AI-specific revenue against AI-specific capital spend. A widening gap between the two is the earliest visible sign of the slower-demand scenario. Capacity growing faster than attributed revenue shows up well before any project gets canceled outright.
The Practical Read on AI Infrastructure
The AI compute buildout is a power procurement story in 2026, not a semiconductor story. Grid interconnection and transformer lead times set the pace, and neither responds quickly to money.
Whether any of this pays back depends on usage that has not fully arrived. The gap between deployed capability and measured economic effect is the subject of our analysis of where AI gains do and do not show up in the data.
Watch regional load growth and utilization rather than headline capital commitments. Emergent Wire tracks the buildout on those terms because announced spending is a plan, and energized capacity is a fact.
Emergent Wire covers AI models, capabilities, and the industry building them.