Skip to content

The Compute Bill Comes Due: Why AI Roadmaps Now Have a Capacity Ceiling

Published: at 01:35 AMSuggest Changes

The platform director brought three AI use cases to a quarterly review, a signed budget and a delivery plan. He left with a fourth item he had not planned for: a capacity problem. His regional cloud provider could commit accelerated compute for one workload, could offer a second only at a materially higher price, and had no firm date for the third. The extra capacity sat behind a grid connection that would not clear for years.

That conversation is spreading through executive committees because the binding constraint on enterprise AI has moved. For three years the limiting factor was model capability — whether a system could reason well enough to be trusted with real work. For a growing set of tasks that question is largely settled. What remains open is whether the organisation can obtain the electricity, the data-centre space and the regional capacity to run those tasks at a price that survives contact with a P&L.

The compute bill is coming due, and most AI roadmaps were written as if it never would. Capacity is now the ceiling on AI ambition. It belongs in the same portfolio conversation as every other scarce resource, rather than in a procurement annex. Leaders who keep treating it as a line item will find their roadmaps capped by physics they never budgeted for.

The ceiling is physical, and it is close

The numbers are no longer speculative. The International Energy Agency estimated that data centres consumed around 415 terawatt-hours in 2024, roughly 1.5% of global electricity, and projected that figure would more than double to about 945 TWh by 2030. In the United States, data centres account for nearly half of electricity demand growth to 2030 — more electricity than aluminium, steel, cement and chemicals combined. [IEA, April 2025]

One year later the IEA updated the picture and found the near-term story tightening rather than easing. Data-centre electricity demand grew 17% in 2025, while consumption at AI-focused facilities rose by half. Bottlenecks across power equipment, grid connections and chip supply were pulling the more aggressive scenarios off the table, even as investment and project pipelines boomed. [IEA, April 2026]

Gartner has been blunter. Its analysts now describe AI scaling as constrained by power availability, with grid access the limiting factor on where capacity can grow. The firm expects data-centre electricity consumption to rise 26% in 2026 and installed power demand to approach 290 gigawatts by 2030. [Gartner, June 2026]

The direction of travel is unambiguous. The scarce input in enterprise AI is no longer intelligence. It is delivered, connected, affordable power and the capacity it feeds.

Intelligence got cheaper; served capacity did not

Something important happened while everyone watched model benchmarks. Frontier model quality rose, and the price per token fell, faster than the physical supply chain could respond. Intelligence became the abundant input in the stack. Served capacity — the combination of chips, power, cooling, grid connection and a region where all of it exists — became the scarce one.

That inversion rewrites the economics of an AI roadmap. A model that is twice as good at half the price is a welcome improvement, but it does not add a single megawatt of capacity. A roadmap that assumes unlimited capacity at falling unit prices rests on an assumption the IEA’s 2026 update has already started to unwind.

The right unit of account follows from this. Token spend and API invoices measure an input. What a CFO can act on is cost per completed task — the fully loaded cost of one successfully finished unit of work, including retries, tool calls, retrieval, monitoring and human review. Capacity planning and unit economics are the same conversation seen from two ends: how much compute each completed task consumes, and how many completed tasks the business actually needs.

An organisation that can state its cost per completed task can rank its workloads. One that cannot is managing a budget it cannot defend.

Forecast compute the way you forecast headcount

Most AI plans are expressed in use cases and benefits. They need to be expressed in demand. The translation borrows from workforce and infrastructure planning.

Start with a task inventory rather than a use-case list. For each production workload, estimate the number of completed tasks per period, the compute each task consumes, and the peak concurrency it requires. Aggregate those into a demand forecast expressed in the units your providers and your grid actually sell: accelerator hours, megawatts and a date by which the capacity must be live.

Then confront lead times. The mismatch here is structural. A data centre can be built and connected to the internet in one to two years, while new transmission to serve it takes four or more years to plan, permit and build. [EPRI, May 2024] Those timelines are set by permitting, equipment supply and grid studies, none of which a procurement team can shorten by negotiating harder.

The queue data confirms the arithmetic. At the end of 2025, around 8,200 projects representing more than 1,300 gigawatts of generation were actively seeking grid interconnection in the United States, and the median time from an interconnection request to commercial operation exceeded five years for projects completed in 2025. Most proposed capacity is never built: only 13% of capacity that requested interconnection between 2000 and 2020 had reached commercial operation by the end of 2025. [LBNL, June 2026]

Two consequences follow. Capacity decisions made today land in a window three to five years out, so a roadmap that assumes next-quarter elasticity is fiction. And the plan must separate workloads that can tolerate that delay from those that cannot.

Rank workloads by value per unit of compute

Once demand is in business terms, prioritisation becomes a portfolio exercise. Rank every workload by the value it creates per unit of compute it consumes, and be honest about the denominator.

The ranking exposes uncomfortable answers. A high-volume assistant answering thousands of trivial queries may show an impressive cost per token and a poor return once each completed task is costed properly. A low-volume agent that closes a compliance gap or prevents a costly error may justify compute several times its footprint. Value per completed task sets the sort order, ahead of volume or novelty.

This is where a capacity ceiling becomes useful rather than merely painful. It forces a portfolio to name what it will stop funding. The candidates are recognisable:

Each of those consumes capacity that the workloads which matter are competing for. Retiring one funds another without a new budget line.

Design for the region, not the global average

Global capacity statistics conceal the constraint that actually bites: availability is local. The IEA found that nearly half of US data-centre capacity sits in just five regional clusters, which is precisely why local grids, rather than national averages, determine what a given organisation can deploy. [IEA, April 2025]

Regional design has three dimensions. Data residency and sovereignty rules decide where a workload may run at all. Latency decides where it should run for an acceptable user experience — inference close to users, training where power is abundant and cheap. Flexibility decides what can move: deferrable batch jobs, non-urgent analytics and model evaluation can follow capacity and price, while interactive customer-facing workloads cannot.

The design question for an architect therefore shifts from which cloud to select to which region, for which workload, under which constraint, and with what fallback when that region is full. A roadmap with a single regional dependency is a roadmap with a single point of failure measured in megawatts.

Capacity is a portfolio with one owner

Most organisations have no single owner of AI capacity. Procurement buys compute, platform teams consume it, product teams request it, and finance reconciles the bill. Nobody holds the whole picture, so nobody can allocate it deliberately.

Fix that with the same discipline applied to any scarce resource. Name one accountable owner for the capacity portfolio. Give them a forecast, a budget and the authority to allocate and reallocate. Price capacity internally so that consuming teams see its real cost, and review allocation on a fixed cadence alongside the workload portfolio.

The decision rights are straightforward. The capacity owner proposes allocation and holds the headroom policy. Business sponsors justify the value per completed task and accept a capacity ceiling for their workload. Security and risk set the residency and control floor that no allocation may breach. Finance validates the unit economics.

Set one standing rule that prevents most of the drift: no new production workload receives a capacity commitment without a forecast demand, a named sponsor and a stated condition under which its allocation is returned. Capacity, once allocated, is otherwise treated as permanent by default — and permanent allocation is how a portfolio loses its ability to fund the next thing.

The leadership test

Ask your platform and finance leads three questions and insist on a number for each. What capacity demand does the current AI roadmap imply, by quarter, in accelerator hours or megawatts? Which named workloads hold the committed capacity, and what value per completed task does each return? Which workloads will you stop funding, or retire, to make room for the next one?

If those answers take more than a week to assemble, or if the roadmap implies capacity that no provider or grid connection can deliver on schedule, you do not have a capacity plan. You have a wish list with a delivery date attached. The ceiling is real, it is close, and the organisations that price it now will be the ones still shipping when the grid runs out of headroom.

Sources


Previous Post
Internal Audit Has to Assure Systems That Change Themselves
Next Post
You Contracted for a Capability, Not a Change Process