Have you noticed how the AI buildout conversation keeps circling the same giant numbers? Gigawatts. Multi-hundred-megawatt campuses. Decade-long offtakes. I have, and I keep coming back to a quieter shift that feels more practical than the headline race. The labs that trained the models everyone talks about are now shopping for much smaller slices of power. Think 20 to 30 megawatts. Not glamorous. Fast.
Why Smaller AI Data Center Deals Suddenly Matter
The story is simple on the surface and messy underneath. Demand for usable compute is outrunning the pace at which enormous sites can be permitted, powered, and filled with chips. When a campus needs new substations, long transmission work, and years of community review, the calendar becomes the real bottleneck. Smaller allocations at sites that already have power look less heroic. They also get workloads online sooner.
I have found that markets often over-index on the biggest announced number and under-index on time-to-first-token. That is the unglamorous metric. How quickly can a cluster serve users? How quickly can a lab add inference capacity without waiting for a single mega-campus to finish? In my experience, operators who can answer those questions win the next round of conversations, even if the megawatt count looks modest on a slide.
Both major frontier labs have spent the past year locking in huge infrastructure commitments. That part is not new. What is new is parallel hunting for compact deployments in places where land, cooling, and grid interconnection already exist. Conversations have reportedly ranged across the United Kingdom, the Nordics, and parts of the United States. The pattern is diversification, not a retreat from scale.
Different workloads need different infrastructure, so conversations happen with a range of partners and get judged on requirements, performance, reliability, timing and cost.
That is the official tone, and it tracks. Training a frontier model still loves dense, tightly coupled clusters. Serving that model to millions of daily sessions does not always need the same topology. Inference can be split. Requests can land on separate smaller clusters. Geography becomes a feature if latency, redundancy, and local power prices cooperate.
Speed To Usable Capacity Beats Bragging Rights
Analysts who watch wholesale compute keep repeating a phrase that should be taped to every planning whiteboard: speed to usable capacity. Securing a few megawatts at an existing powered site can beat waiting for a much larger block in one location. If workloads can run across separate sites, a handful of compact deals can add up to serious capacity without a single ribbon-cutting ceremony.
Perhaps the most interesting aspect is how unromantic this is. Nobody writes poetry about a 25 MW hall on an industrial park. Yet that hall might be live while a flagship campus is still arguing over water use and transformer lead times. I would rather have boring capacity that ships tokens than a spectacular rendering that ships press releases.
- Existing interconnection often beats a greenfield promise.
- Shorter construction cycles reduce execution risk.
- Distributed sites can isolate outages and regional shocks.
- Contract structures can be more flexible than decade-long campus deals.
None of this means giant projects disappear. They remain the backbone for the heaviest training runs. The portfolio just gets wider. Labs want a mix: dense training campuses, regional inference nodes, and overflow capacity that can be rented when a product spike hits.
Training Clusters Versus Inference Footprints
Training a large model typically requires many chips working closely together. High-bandwidth interconnects. Tight latency. A room that behaves like one machine. Inference is often a different animal. Many inference workloads can serve separate requests across multiple smaller clusters. That opens more locations, including markets that could never host a gigawatt campus.
The mix of work inside data centers is expected to tilt. In the mid-2020s, training still commanded a meaningful share of specialized AI capacity, while inference sat lower. Industry real-estate research has projected inference rising sharply later in the decade, eventually using a far larger slice of total AI-oriented capacity than training. If that path holds, site strategy has to follow the workload, not the other way around.
I keep a simple mental model. Training is a factory. Inference is a retail network. You still need the factory. You also need stores close to customers, with enough inventory to handle a Friday night rush. Smaller halls are those stores. They will not replace the factory. They keep the product on the shelf.
| Workload | Typical cluster style | Site preference |
| Frontier training | Dense, tightly coupled | Large powered campuses |
| Fine-tuning | Medium clusters | Flexible regional halls |
| Production inference | Many smaller pods | Distributed live sites |
| Burst serving | Overflow rental | Neocloud and colo mix |
That table is a sketch, not a law. Some inference still wants fat pipes and co-located memory. Some training can be staged. The point is optionality. Labs that only know how to buy one kind of building will get stuck when the product mix shifts.
Why Europe And The Nordics Keep Coming Up
Europe is not an easy place to drop a gigawatt box. Land is tight. Power queues are long. Local opposition has grown louder. Cooling water is politically sensitive in more than one region. So why hunt there at all? Because pockets of ready power still exist, especially where industrial grids, hydro, or surplus capacity can support a smaller hall without a national drama.
The Nordics have a reputation for cooler ambient temperatures, relatively clean electricity, and operators who already speak the language of high-density compute. The United Kingdom has demand sitting next to constrained supply, which makes any ready megawatt valuable. Smaller deals can slip into that gap. They do not solve the continent’s power shortage. They harvest the leftover sockets.
In the United States the constraint looks different but rhymes. Communities push back on mega-campuses. Transmission upgrades slip. Transformer deliveries take forever. A 20 to 30 MW expansion at a site that already has a substation can feel almost polite by comparison. Polite still counts when the alternative is a three-year wait.
Community Pushback And The Politics Of Power
Huge data center projects face more local resistance than they did two years ago. Residents worry about noise, water, traffic, tax abatements, and electricity prices. Some of those worries are fair. Some are exaggerated. All of them slow calendars. The industry has not always done a good job explaining tradeoffs. That vacuum fills with rumor.
Smaller footprints do not erase politics. They can lower the temperature. A compact hall next to existing industry looks less like a new city of servers. It still draws power. It still needs cooling. It is simply easier to site than a campus that wants its own river and a dedicated transmission line.
I’ve found that the smartest operators treat community work as part of the critical path, not a press afterthought. If you need neighbors to accept night-time generators or a new feeder, you start that conversation before the lease is signed. Sounds obvious. Plenty of projects still skip it and then act surprised when hearings drag.
How The Commercial Structure Actually Works
Frontier labs rarely want to own every brick. They rent capacity from data center operators and so-called neoclouds that assemble power, cooling, networking, and accelerators into a billable service. Large-scale, long-term agreements remain the default for flagship training. Smaller deals can be shorter, staged, or tied to specific product launches.
One lab has already been linked to a very large cloud-style commitment measured in hundreds of megawatts at a single development. Another has publicly described infrastructure programs that started with multi-gigawatt ambitions and then added more regional capacity. Those headline numbers still matter. The parallel 20 to 30 MW conversations matter because they fill the months in between.
- Identify a powered shell or live hall with spare density.
- Match chip generation, networking, and cooling to the workload.
- Negotiate term, take-or-pay, and expansion rights.
- Stage delivery so inference can go live before the last rack lands.
- Operate across sites as one logical pool where the software allows it.
That sequence looks tidy. Reality is messier. Chip allocations slip. Liquid cooling kits arrive late. A utility changes an interconnection date. The advantage of a smaller deal is that one slip does not freeze an entire national plan. You can keep shipping on site A while site B waits for a transformer.
Neoclouds, Colos, And The Middle Layer
A new middle layer of compute suppliers has grown up around the AI boom. Some started as energy or crypto-adjacent operators and pivoted into GPU hosting. Some are traditional colocation firms that learned high-density cooling in a hurry. A few are building both giant complexes and a second line of smaller, faster halls.
Why the second line? Because large builds face delays across several U.S. markets. Faster, cheaper halls can be stood up on sites that already have power. That is not a theory. It is a capital allocation choice. Investors have been willing to fund those operators at high valuations because contracted demand looks real and the queue for chips still feels tight.
Is every neocloud a durable business? Of course not. Some will overbuild. Some will mis-time chip generations. Some will sign customers who later concentrate spend with a hyperscaler. The ones that survive will probably be good at two unsexy skills: delivering power on the date promised, and keeping clusters from turning into expensive space heaters.
What 20 To 30 Megawatts Actually Buys
People hear “small” and imagine a closet. That is the wrong picture. Twenty to thirty megawatts of IT load is still a serious industrial facility. Depending on density and generation, it can host a large number of accelerators. It can support a meaningful inference region. It can also host specialized fine-tuning or evaluation clusters that do not need to sit next to the main training fabric.
Do the rough math in your head. A modern high-density hall might pull tens of kilowatts per rack. Liquid cooling changes the ceiling. Networking and storage eat some of the budget. Even after overhead, a compact site is not a toy. It is a node. String enough nodes together and you have a network.
Portfolio sketch, not a forecast: Flagship training campuses — multi-hundred MW to GW class Regional inference nodes — 20 to 50 MW class Burst and overflow — rented by the cluster Edge-ish serving — smaller still, latency driven
I like that sketch because it refuses the false choice between “one giant temple” and “a thousand tiny boxes.” Real systems are layered. Finance teams should underwrite the layer, not only the temple.
Chip Supply, Power Price, And Timing Risk
Even a ready building is useless without accelerators. The chip cycle still governs everything. A hall that goes live six months before boards arrive is a very expensive warehouse. That is why timing conversations now include vendors, integrators, and utilities in the same room. The bottleneck moves. Last year it was GPUs. This year it might be substations. Next year it might be memory or networking silicon.
Power price is the other quiet killer. A cheap lease on a pricey grid is not cheap. Labs look at effective cost per token, not just rent per kilowatt. Regions with abundant low-carbon power still attract interest for that reason, provided interconnection is real and not a slide-deck promise.
There is also a human timing risk. Teams that design clusters want the latest generation. Product teams want capacity yesterday. Finance wants utilization. Those three clocks almost never agree. Smaller deals can absorb some of that disagreement because you can refresh a 25 MW hall without re-planning a national campus.
Investors Should Watch Utilization, Not Just Announcements
Markets love a signed memorandum. Utilization is duller and more honest. If a lab fills compact sites quickly, that is a signal that inference demand is real and geographically spread. If those sites sit half empty while giant campuses keep getting announced, the strategy may be more about optionality than immediate load.
Watch three things. First, the mix of training versus serving in public comments. Second, the cadence of smaller offtakes, even when they are not branded. Third, local grid news in the same counties as new halls. Power stories and compute stories are now the same story wearing different jackets.
- Contracted megawatts versus energized megawatts
- Chip generation actually installed, not just reserved
- Latency targets by product region
- Community and permitting timelines
- Effective power cost after incentives
None of those fit neatly on a earnings-call slide. They still decide who can ship features when a rival model drops.
The Software Layer Makes Small Sites Viable
Hardware people sometimes forget that software decides whether a distributed fleet works. If orchestration, routing, and model sharding can treat separate halls as one pool, compact sites become much more attractive. If every cluster is an island with its own quirks, operations teams will fight the map.
That is why the inference shift is not only a real-estate story. It is a systems story. Better schedulers, better caching, better request batching, and smarter placement of model replicas all raise the value of a 30 MW node in a second country. Without that software, you just own more buildings.
I’ve sat through enough architecture reviews to know the temptation: design for the beautiful single fabric, then bolt on regions later. The labs that invert that habit will have an easier time using leftover power in awkward places. Awkward places are where spare megawatts now live.
Risks That Do Not Show Up In The Press Release
Fragmentation is a real cost. More sites mean more vendors, more network paths, more compliance regimes, more night-shift pagers. Security teams hate sprawl. So do finance teams who suddenly have twenty small invoices instead of two large ones. There is a point where “agile capacity” becomes “unmanageable capacity.”
Another risk is stranded generation. If a lab bets on a chip family that ages out faster than the lease, a compact hall can still become awkward inventory. Smaller does not automatically mean flexible. Contract language does.
Water, noise, and diesel backup remain local flashpoints even at modest scale. Do not assume a 25 MW site is invisible. It is simply less cinematic than a gigawatt rendering. Neighbors can still organize. Utilities can still say no. Treat every site as a political object, because it is.
Securing a few megawatts at an existing powered site can be more practical than waiting for a much larger block in one location.
What This Means For The Next Two Years
Expect a barbell. On one end, a handful of enormous training campuses that keep growing in public view. On the other, a quieter string of compact halls that turn inference into a distributed service. The middle will be messy: delayed campuses, half-finished shells, and operators who promised dates they cannot keep.
Also expect more study of purpose-built smaller facilities designed around distributed inference rather than around a single training run. Chip vendors have already signaled interest in that design space. Energy-first operators are experimenting with faster builds. Traditional landlords are rewriting density rules. The building type is still being invented in public.
If inference really does claim a rising share of specialized capacity through the late 2020s, the map of “important” data center markets will widen. Not every city becomes a training capital. More cities can become serving nodes. That is a different real-estate thesis and a different power thesis. Investors who only underwrite the temple will miss the network.
A Practical Checklist For Operators And Buyers
If you run sites, stop selling only scale. Sell energized dates. Sell cooling that matches the next two chip generations. Sell expansion rights that do not require a new act of parliament. If you buy capacity, ask for the ugly calendar, not the glossy one. Who owns the interconnection queue position? What happens if boards slip by a quarter? Can the cluster join an existing routing domain on week one?
- Prefer powered shells over artistic master plans.
- Write contracts that allow partial go-live.
- Measure cost per served request, not only cost per megawatt.
- Budget people and process for multi-site operations.
- Keep a training core; do not pretend every watt is interchangeable.
That last point matters. Distributed inference is not a license to abandon dense training fabric. The labs still need rooms where thousands of accelerators behave like one instrument. The new work is holding both ideas in the same plan without letting either one become religion.
The Human Tempo Behind The Megawatts
There is a temptation to treat this as pure machinery. It is also a staffing story. You need electricians, liquid-cooling techs, network engineers, and night operators who can handle a compact hall in a market that never hosted AI density before. Training those teams takes time. Importing them takes visas and housing. The calendar again.
Product managers live on a different clock. They want a new region because a customer in that region is waiting. They do not care that a transformer is sitting on a dock. Smaller deals exist to reconcile those clocks a little. Not perfectly. A little. In this industry, a little time saved is a feature launch that actually happens.
I do not buy the idea that compact sites are a fad. They look like the obvious response to a world where power is lumpy, permits are slow, and serving work is growing faster than the appetite for another five-year groundbreaking photo. The giants will keep building giants. They will also keep knocking on doors that already have power. That combination is the real race.
So when the next announcement boasts another enormous campus, read the fine print and then look sideways. The interesting capacity might be a quieter 25 MW hall that starts serving next quarter. Less cinema. More tokens. That is usually how infrastructure actually moves.