25 Aug 2026
Signal Headquarters
Vol. I
No. 237

AI token costs per knowledge worker increase with each model improvement, as users delegate more tasks at higher model quality.

The case

The most valuable tokens to spend are on agents where ROI per token can be directly measured, rather than on generic AI adoption.

“Tier three tokens the most valuable are this agents where you can get the ROI of each specific token and I can do that now.”
Carlos García · 10 Aug 2026

Decagon's token usage per conversation has increased over time because more model calls are added to improve quality, contrary to the assumption that optimization reduces token usage.

“Actually over time, the number of tokens we're using per conversation has gone up because we're actually doing more model calls to make the quality better.”
Jesse Zhang · 31 Jul 2026

Modern models such as o3/o4 (referred to as '5.5') can think for weeks when scaffolded before performance plateaus, making plateau-based evaluation impractical.

“What we're seeing today with the modern models is that 5.5 and other models can think for if you scaffold them reasonably well, can think for weeks even before having performance plateau.”
Noam Brown · 26 Jun 2026

A common emerging workflow is to use a high-quality, expensive model such as Claude Opus or GPT to generate a taste file for a project, then use cheaper models for all subsequent development work on that project.

“A lot of people what they're doing is they're building one project with a really high quality LLM like Opus or GPT They're building a taste file and then you know, super cheap models to continuously build on that more, you know, project with that taste file.”
Ahmad Awais · 6 Jun 2026

AI token spending will consume one-third of engineering salaries and drive a $4 trillion market cap outcome within two years.

“It's going to eat one/ird of engineering salaries and that's going to get you to 4 trillion by two years.”
Jason Calacanis · 4 Jun 2026

Anthropic has become a premium product since the end of last year, priced at twice the cost of its competitor.

“It has become a premium product since the end of last year. It is twice the cost of its competitor, right?”
Jason Lemkin · 28 May 2026

The pushback

Within the next year, teams will be able to dial in compute per task, cutting token costs by over 90%.

“I think you'll get to a place probably over the next year where you can like really dial in hey how much compute do I want to spend on this because I have certain like cost considerations and certain latency considerations and get to like the exact optimal amount of cost. and so if you do that like your token costs go down 90% plus.”
Nathan Labenz · 22 Aug 2026

An Exo agent autonomously rearchitected its own Discord adapter at runtime, scoping context to specific conversations and threads, and reduced costs by 96%.

“It went and rearchitected its own Discord adapter at runtime. Made changes, observed them, tested them to really scope down the context that was included in each of the LM calls as they were built by scoping it to certain conversations and certain threads only rather than pulling messages from across different threads in Discord in a way that drove down the cost to I think like it was like a 96 decrease 96% decrease.”
Alex Krentsel · 15 Aug 2026

Silico's $1,000/month price will decrease over time as Goodfire develops more token-efficient agent methods.

“My hope is we're starting out with $1,000 a month subscription. We'll be able to bring that down over time because we're able to come up with more and more clever ways.”
Dan Balsam · 8 Aug 2026

Claude Opus generates 70% fewer tokens for the same question compared to earlier model versions.

“Claude is even on Opus is generating 70% less tokens for the exact same question.”
Gavin Baker · 20 May 2026

Tasklet opted not to ship Claude Opus 4.7 as the default recommended model due to increased cost and limited benefit for iterative knowledge work.

“We actually opted not to ship 47 as like a like a default recommended model. We are going to ship it but as sort of an advanced user option if people want to and we're going to like note that hey this is actually like costs a lot more.”
Andrew Lee · 15 May 2026

Topics

AI AdoptionAI AgentsAI Model Economics

Signal Headquarters · compiled from attributed public discussion. Last updated 2026-08-22.