Inference cluster scale is rapidly expanding from single-digit chip configurations to thousands or tens of thousands of chips.
The case
AI labs are already training models across multiple data centers, treating entire data centers as a single machine.
“Some of my friends at labs right now are training models across data centers. It's not even across nodes anymore. It's not even racks.”Nathan Labenz · 22 Aug 2026
Google has already made a deal to sell approximately 20% of its TPUs to Anthropic.
“Google already made a deal to sell like 20% of their TPUs to anthropic.”Patrick O'Shaughnessy · 18 Aug 2026
A trillion dollars of capex has already been committed for the next year across major cloud and AI players.
“There's a trillion dollars of capex that has been committed for the next one year across all these people broadly speaking.”Harry Stebbings · 6 Aug 2026
vLLM is currently running on half a million GPUs at any moment.
“The open- source inference engine now running on half a million GPUs at any moment.”Simon Mo · 6 Aug 2026
Inference providers like Fireworks will become far more capex-intensive as they vertically integrate into owning data centers.
“These are going to become way more capex intensive businesses.”Rory O'Driscoll · 23 Jul 2026
Token count processed per day will increase 20x to 100x by the end of next year.
“Anywhere ranging from 20 to 100x could be possible.”Lin Qiao · 20 Jul 2026
The pushback
GPUs cannot effectively use tensor parallelism beyond 4-8 chips due to communication overhead from inability to overlap communication and computation.
“One of the limits of GPUs is because they don't effectively overlap communication and they have a hard time using tensor parallelism beyond four or eight.”Kunle Olukotun · 9 Jul 2026
Topics
Signal Headquarters · compiled from attributed public discussion. Last updated 2026-08-22.