Inference clusters are scaling from single-digit chips to tens of thousands, and the infrastructure world has not caught up
The mental model of inference as an 8-chip problem is already obsolete. Token demand is set to grow 20x to 100x by end of next year, and the hardware required to serve it is evolving faster than most operators can track.
The working assumption across much of the industry is that inference is a contained, predictable infrastructure problem. Gavin Uberti, co-founder of Etched, thinks that assumption is already wrong. He describes inference clusters moving very quickly from today’s 8-chip configurations, or NVL72 scaleup domains, to thousands of chips and then tens of thousands. The scale that was considered exotic for training workloads a few years ago is becoming the baseline expectation for serving models at production demand.
The demand numbers Lin Qiao, CEO of Fireworks AI, puts on this trajectory are not incremental. She projects that daily token throughput could grow anywhere from 20x to 100x by the end of next year. That range is wide enough to reflect genuine uncertainty about adoption pace, but even the lower bound implies infrastructure requirements that current cluster architectures are not designed to absorb. The gap between where inference capacity sits today and where demand is heading is not a rounding error.
Uberti adds a second dimension to the argument: interconnect. He notes that chip-to-chip latency could theoretically be reduced to as low as 2 nanoseconds, implying enormous headroom relative to what current systems achieve. That figure points to a physical frontier that has barely been approached, which means the scaling story is not purely about adding chips. The quality and speed of the connections between them will determine whether large clusters can actually coordinate at the latencies that useful inference requires.
After three years, if every year there's three hardware skill, after three years there are nine hardware skill in between. Do you still want to go back to nine generation older hardware running three years old model on that? That's questionable. Lin Qiao
Hardware lifecycle compression is a separate pressure that Lin Qiao describes in terms that warrant attention. She estimates the industry is producing approximately three new hardware SKUs per year. At that pace, hardware that is three years old sits nine generations behind current state of the art, a gap she calls questionable for running modern models. The implication is not simply that older hardware underperforms. It is that depreciation cycles, procurement planning, and data center investment horizons are all calibrated to a refresh rate that no longer matches the actual tempo of the market.
Dan Biderman, CEO of Answer.AI, adds a data-layer argument that compounds the compute picture. He projects that many AI-native companies will be managing trillions of tokens of internal proprietary data within 18 months. His framing is careful: he calls it an exaggerated-sounding figure that he does not think is an impossibility for companies that are genuinely AI-native. If that projection holds, the inference compute required to reason over those knowledge workspaces at useful latency will stack on top of the external-demand growth Qiao describes, pushing cluster requirements further still.
The external infrastructure buildout is already reflecting this direction. Public reporting from China Daily notes that China’s Dawning 8000 supercluster in Zhengzhou, the country’s first capable of supporting over 100,000 domestically developed computing cards, went live in July. Separately, analysis from Squared Tech observes that 100,000-GPU systems signal that AI leadership increasingly depends on operating an entire computing stack, not merely buying leading chips. Both data points suggest that the cluster-scale transition Uberti and Qiao describe is not a future forecast waiting for confirmation. It is already being built into capital allocation decisions at national scale.
What connects these claims is a consistent direction of travel: inference is becoming a large-scale distributed systems problem at a speed that most procurement, networking, and cooling infrastructure was not designed to handle. The 8-chip mental model was never wrong for the workloads that existed when it formed. The workloads have changed. The clusters are following. The question now is whether the rest of the stack, from interconnect design to hardware refresh economics to data management, moves fast enough to serve the demand that Qiao’s 20x-to-100x range already implies is coming.