21 Aug 2026
Signal Headquarters
Vol. I
No. 226
· · 3 min read

Cheaper AI tokens do not shrink AI bills, they multiply them

Every time frontier AI models get cheaper, consumption rises faster than anyone predicts. The pattern has a name, it has historical precedent, and it is now the central fact of AI economics.

Krishna Rao puts the dynamic plainly: when Anthropic lowered the price of its Opus models, consumption rose far more than anyone expected. That is Jevons paradox operating in real time, and it is now the defining force in AI economics.

The pattern holds at the company level too. Andrew Feldman, chief executive of Cerebras, reports that eight months ago his firm was not spending $1,000 per engineer on tokens. The figure now sits at $25,000 to $30,000. That is not a gradual drift; it is a structural shift in what a professional workstation costs. Cat Wu observes the same dynamic at the product level: as models improve, people delegate far more tasks to them, spend more hours inside them, and the token cost per engineer or knowledge worker rises with every model jump or substantial product improvement.

The supply-side math makes clear why consumption keeps climbing even as per-token prices fall. Harry Stebbings describes a compounding improvement: raw chip performance roughly triples every 18 months, and additional optimizations, including quantization and other inference techniques, add another roughly 3x on top of that. The result, he argues, is approximately a 10x improvement in tokens per unit of money every couple of years. Dylan Patel offers a more specific data point: DeepSeek, he notes, achieved GPT-4-level performance at one six-hundredth the cost. Jason Calacanis draws the structural conclusion that any frontier model becomes five to ten times cheaper within a year of being superseded.

History offers a useful frame for what happens when a technology gets radically cheaper and more convenient. Dara Khosrowshahi, Uber’s chief executive, describes what happened when ride-hailing undercut the taxi market: analysts looked at the size of the black car marketplace and the taxi industry and calculated the ceiling from there. What they missed was that radically improving convenience or cutting cost expands the market beyond any original estimate. As Khosrowshahi puts it, “the company today is a result of Jevons.” Harry Stebbings adds a concrete data point from inside Uber itself, citing a disclosure from the company’s chief operating officer that Uber spent an entire year’s worth of Anthropic credits in just four months. The AI token market is following the same logic, and the consumption figures are already outpacing the forecasts built around falling prices.

As the models get better people delegate far more tasks to it and they spend a lot more hours in. We do see the token cost per engineer or like per any knowledge worker increase every time that there's a model jump or like a substantial product improvement. Cat Wu

The demand projections that flow from this logic are large enough to invite skepticism, but they come from multiple directions. Gavin Baker argues the shift to usage-based pricing is why OpenAI and Anthropic will together exceed $200 billion in annual recurring revenue this year. Patel projects the economy will spend $100 billion on a single tier of frontier model by the end of the year, up from roughly $40 billion currently. Stebbings cites a forecast that agents alone will push token consumption up 24 times by 2030. Each of these figures carries meaningful uncertainty, but they share a direction.

Feldman’s arithmetic on software engineering illustrates how the numbers can become very large very quickly. With 47 million software engineers in the world, he calculates, software engineering token use alone represents a $5 trillion market. The framework is the point: at scale, even modest per-seat consumption figures produce aggregate numbers that dwarf current AI revenue.

The evidence from individual products reinforces the macro picture. Jesse Zhang notes that token usage per conversation has risen over time because more model calls are added to improve quality, contrary to the assumption that optimization reduces total consumption. Roman Chernin reports that the same week a major AI infrastructure company’s stock dropped 40 percent in early 2025, his firm had its best sales week ever, suggesting that cost-reduction announcements accelerate adoption rather than depress it. Jesse Genet adds a consumer dimension, estimating that a family genuinely using AI could spend somewhere between $2 and $500 per month on tokens, and predicting that if monthly household costs approach $400, demand for local models will spike, driven by cost rather than privacy concerns.

What the evidence describes, taken together, is a market where falling prices and rising consumption are not in tension. They are the same phenomenon. Every efficiency gain gets consumed by expanded use, every cheaper tier attracts more volume, and every model improvement causes users to delegate tasks they would not have attempted at the prior cost. The companies and institutions that built AI budgets around the assumption that cheaper tokens mean smaller bills are, by this logic, building on a misreading of how the market actually works.

The Editor, for the readers of Signal Headquarters

AI CostAI EconomicsAI InferenceAI Model UsageAI PricingAI RevenueJevons Paradox



From the Archive