Memory, not compute, is the binding constraint on AI infrastructure buildout
Hyperscalers are committing half their capital spending to memory, supply is running 30 percent short of demand, and meaningful new capacity will not arrive until 2028. The bottleneck is not chips in the abstract. It is specifically memory, and the consequences are already visible in prices, margins, and the architecture of the models being built.
Hyperscalers are spending half their capital budgets on memory this year. That single figure, supplied by Dwarkesh Patel, reframes the AI infrastructure story. The dominant narrative has centered on compute, on GPU clusters and chip shortages and Nvidia’s order books. The capital allocation data points somewhere else entirely.
Martin Casado adds the broader financial context: Microsoft, Meta, and Google are each on track to spend over 50 percent of their total revenue on capital expenditure. That is an extraordinary commitment by any standard. When Patel’s figure is placed alongside it, the implication is that memory is consuming a share of corporate revenue that would have seemed implausible to infrastructure planners a few years ago. Public analysis from TrendForce puts memory’s share of cloud service provider capital spending at 47 percent in 2026, climbing toward 68 percent in 2027, a trajectory that corroborates the direction if not every specific number.
The supply side of that equation is under acute stress. Patrick O’Shaughnessy estimates the market is already 30 percent short of demand. Dylan Patel warns that true incremental supply from new capacity investments will not arrive until 2028, which he calls “a very unique thing.” The gap between current demand and future supply is not a near-term blip to be smoothed by the next product cycle. The constraint runs for years. The downstream effects are already materializing: Caitlin Kalinowski expects prices to roughly double, and is actively advising startups to pre-buy memory and hold enough stock to ride out price spikes. Jake Cooper observes that servers have actually appreciated in value as RAM prices rise, an inversion of normal hardware depreciation that signals how dislocated the market has become.
We are not in a computebound world. We are in a memory and network and communication or bandwidthbound world. Nathan Labenz
The margin structure of memory producers reflects the same pressure. Nathan Labenz notes that memory chip makers are now operating at 80 to 90 percent gross margins, above even Nvidia’s roughly 70 percent. Andrew Feldman points specifically to Micron, describing gross margins of 80 to 85 percent. When a commodity supplier posts margins at that level, the market is telling you something unambiguous about the relationship between supply and demand.
What makes the bottleneck technically interesting is that the binding constraint is not simply how much memory exists. It is how fast memory can move data. Reiner Pope identifies memory bandwidth, not raw compute or total memory capacity, as the operative limit on inference. Kunle Olukotun frames the same problem from a different angle: as models grow, running inference requires moving weights and what practitioners call the KV cache into compute units, and that is fundamentally a data movement problem, not a computation problem. Labenz states the consequence plainly: the AI hardware constraint is now memory and bandwidth rather than compute.
Pope offers the most concrete evidence that this is not a theoretical concern. Context lengths in frontier models climbed steeply from around 8,000 tokens to 100,000 to 200,000 tokens. Then, for roughly the past year or two, they stopped. Pope argues that the plateau is not a modeling choice. It is a cost ceiling imposed by memory bandwidth. Pushing context lengths dramatically further would be cost-prohibitive, not because of compute expense but because of memory bandwidth cost. A stagnation in one of the most visible capability metrics of frontier AI, traced directly to a physical supply constraint, is a more concrete consequence of the memory bottleneck than any margin figure.
The price signal has spread beyond memory chips. Yaroslav Azhnyuk reports that optic fiber went from roughly four dollars per kilometer to roughly 32 dollars per kilometer in a matter of months. That is a different component, but the same dynamic: infrastructure demand running far ahead of supply, with prices repricing rapidly to reflect the gap. The memory crunch is not isolated. It is the sharpest expression of a broader constraint on the physical buildout of AI infrastructure, one that capital commitments alone cannot resolve on any timeline shorter than several years.