29 Aug 2026
Signal Headquarters
Vol. I
No. 271
· · 3 min read

Nathan Labenz is betting that per-task compute control will cut AI token costs by 90 percent within a year

The prediction is not that models will get cheaper in the abstract. It is that fine-grained control over compute per task will eliminate the overspend baked into how most teams deploy AI today. The mechanism is concrete enough to be checked.

Nathan Labenz is predicting a sharp, near-term collapse in AI token costs, and the mechanism he names is precise enough to be tested. The bet is not that models will get cheaper in the abstract. It is that teams will gain fine-grained control over how much compute they spend on any individual task, and that doing so will cut costs by more than 90 percent.

The reasoning turns on a distinction that most current AI deployments ignore. Right now, many teams route every request through the same model at the same effort level regardless of what the task actually requires. A query that needs a single lookup gets the same compute as one that requires multi-step reasoning. Labenz argues that once cost and latency considerations can be dialed in precisely, that mismatch disappears. “I think you’ll get to a place probably over the next year where you can really dial in, hey, how much compute do I want to spend on this, because I have certain cost considerations and certain latency considerations, and get to the exact optimal amount of cost,” he said. “If you do that, your token costs go down 90% plus.”

The timeframe he names is roughly a year. That makes this a falsifiable call: either meaningfully granular, per-task compute control arrives and delivers that order of cost reduction at scale, or it does not. There is no version of this claim that can be quietly revised into vagueness once the deadline passes.

I think you'll get to a place probably over the next year where you can really dial in hey how much compute do I want to spend on this because I have certain cost considerations and certain latency considerations and get to the exact optimal amount of cost. and so if you do that your token costs go down 90% plus Nathan Labenz

What the Labenz prediction requires, in operational terms, is tooling that lets teams match compute to actual task complexity rather than defaulting to maximum effort across the board. The 90 percent figure reflects the gap between what low-complexity tasks actually need and what they currently consume. Close that gap consistently across a mixed workload, and the arithmetic works. Whether the tooling to enable that matching arrives and proves reliable at scale within the window Labenz describes is precisely what the next year will settle.

What has not yet arrived is the operational habit to use such control even where it exists. The teams best positioned to capture a 90-plus percent cost reduction are not the ones waiting for a cheaper model. They are the ones that classify their workloads by actual computational demand, set low effort as the default, and escalate only when a task genuinely warrants it. That is a routing decision as much as a model decision, and most organizations have not yet made it.

The current pattern runs in the opposite direction. As model quality rises, users delegate more ambitious tasks, and token costs per knowledge worker climb. That is not irrational behavior. Delegating more is the point. But it means the cost reduction Labenz describes will not arrive automatically. It requires deliberate effort calibration, and the teams that fail to build that discipline may find themselves paying more, not less, as model capability continues to rise.

The 90 percent figure deserves scrutiny precisely because it is the kind of number that sounds like marketing. But the mechanism Labenz cites is concrete: match compute to task requirements instead of applying maximum effort uniformly, and the overspend on low-complexity tasks disappears. Whether that holds at the level of real mixed workloads rather than idealized ones is the question a year of deployment data will answer.

The Editor, for the readers of Signal Headquarters

From the Archive