Inference infrastructure is scaling by orders of magnitude, and the capital is already committed
The standard inference cluster today is an 8-chip box. Within a very short window, that unit is headed toward thousands of chips and then tens of thousands. The capital, the demand, and the engineering rationale are all pointing the same direction at the same time.
Gavin Uberti, who works on inference hardware, describes the current baseline plainly: the standard inference scaleup domain today is an 8-chip cluster, or at most an NVL72 configuration. That baseline, he argues, is about to become obsolete. Very quickly, in his framing, inference clusters will move to thousands of chips and then tens of thousands. The engineering question is whether the interconnect technology can keep pace with that ambition, and Uberti’s answer is that the physical limits permit it: latency between chips, he says, could come down to as little as 2 nanoseconds.
That shift is not purely architectural speculation. Nathan Labenz, who tracks AI infrastructure closely, reports that the training layer has already crossed a comparable threshold. “Some of my friends at labs right now are training models across data centers,” he says. “It’s not even across nodes anymore. It’s not even racks.” The unit of computation has jumped from a rack to an entire facility, and what was once an analogy is now an engineering reality. What training has already done, inference is preparing to follow.
The capital commitments behind this expansion are correspondingly large. Harry Stebbings puts a number on the aggregate: “There’s a trillion dollars of capex that has been committed for the next one year across all these people broadly speaking.” That figure, sourced to committed spending rather than projected ambition, sets a floor on what the industry believes the infrastructure buildout will cost. Rory O’Driscoll adds a structural observation: inference providers are becoming far more capital-intensive businesses as the scale of the infrastructure they must own and operate grows. Not every committed dollar will translate into operating infrastructure on time, but the direction and the magnitude of the bet are not seriously in dispute.
Patrick O’Shaughnessy notes that Google has already made a deal to sell roughly 20 percent of its Tensor Processing Units to Anthropic, an arrangement that illustrates how the allocation of existing compute is shifting even before new capacity comes online. David Friedberg adds a counterweight: one major player has reportedly scaled back a planned compute buildout from approximately 1.4 trillion dollars to roughly 600 billion. That revision is large in absolute terms, but the remaining figure is itself staggering, and it suggests the question is not whether large-scale infrastructure gets built but how large and by whom.
Some of my friends at labs right now are training models across data centers. It's not even across nodes anymore. It's not even racks. Nathan Labenz
On the demand side, Lin Qiao projects that token processing volumes could grow anywhere from 20 to 100 times within the next year. Simon Mo reports that vLLM, the open-source inference engine, is already running on half a million GPUs at any given moment. These are not forward-looking projections so much as current operating conditions pointing toward the scale the industry is building for.
The economics of token costs are pulling in the same direction. Brad Gerstner notes that token costs are doubling every 45 days, a compression rate that makes larger-scale inference economically viable faster than most cost models assumed. Clay Bavor observes that top engineers leaning into AI coding agents are already spending more than $100,000 per year on tokens, and that those engineers estimate productivity gains of between three and 20 times in features shipped. The demand is not hypothetical; it is showing up in enterprise spending patterns now.
Lin Qiao adds a hardware-lifecycle dimension that sharpens the urgency. After three years at roughly three new hardware generations per year, a cluster built today would be nine hardware generations behind current silicon. At that point, running modern models on that older hardware becomes questionable. For operators making infrastructure decisions now, the pace of hardware turnover is itself a forcing function: delay compounds quickly.
What the evidence describes, taken together, is a transition already in motion rather than one that is merely anticipated. The scaleup domain for inference is not inching upward; it is moving by orders of magnitude within a compressed window. Dan Biderman captures a parallel dynamic on the data side, observing that AI-native companies could be managing trillions of tokens of internal proprietary data within 18 months. The infrastructure being assembled today is not sized for current workloads. It is sized for what the next generation of AI-native operations will demand, and the evidence suggests that demand is arriving on roughly the same schedule as the buildout meant to serve it.