Local AI models running on consumer hardware are good enough for ~80% of use cases, and will reach frontier model capability within 18 months.
The case
Within the next year, teams will be able to dial in compute per task, cutting token costs by over 90%.
“I think you'll get to a place probably over the next year where you can like really dial in hey how much compute do I want to spend on this because I have certain like cost considerations and certain latency considerations and get to like the exact optimal amount of cost. and so if you do that like your token costs go down 90% plus.”Nathan Labenz · 22 Aug 2026
Overlap built a fine-tuned video understanding model that significantly outperforms major LLMs at understanding video.
“We have our own fine-tuned video understanding model that is way better at understanding video and the major LLM are not good at understanding video.”Sam Parr · 21 Aug 2026
Training a model close to frontier capability is not technically hard today, and many actors are achieving this not just via distillation.
“It's not hard to train a model that is close to frontier capability.”Dan Balsam · 8 Aug 2026
Open-weight models can offer up to 10 different speed tiers, versus only two (regular and fast) for proprietary model APIs.
“For proprietary model there is regular mode and fast mode and that's only the two switch here. But for openweight when you're running it, every provider can offer potentially even 10 different levels of speed going from like the slowest mode which can be a lot cheaper to 400 tokens per second almost up to 500 in many cases that for some workloads and this is typically 2x or 3x faster than the fast mode out there today.”Simon Mo · 6 Aug 2026
Fine-tuned smaller models can be simultaneously better, cheaper, and faster than frontier models on specific tasks.
“We end up getting all three things. It is better at the toss. It is cheaper and it is faster.”Jesse Zhang · 31 Jul 2026
AI model training costs must be recovered within 12-24 months due to rapid model obsolescence.
“You got to recover it pretty damn quick because it only lasts, you know, 12 24 months before it's obsolete.”Rory O'Driscoll · 23 Jul 2026
The pushback
Current open models do not yet have a complete mapping of the social physics of humanity.
“I don't think the model has yet at least the models that are out in the open has yet learned the complete mapping of social physics of humanity.”Joon Sung Park · 21 Aug 2026
Hardware lifecycles are accelerating at approximately three new SKUs per year, making hardware that is three years old nine generations behind, rendering it questionable for running modern models.
“After three years, if every year there's three hardware skill, after three years there are nine hardware skill in between. Do you still want to go back to nine generation older hardware running three years old model on that? That's questionable.”Lin Qiao · 20 Jul 2026
Frontier AI models will spend the next one to two years building memory and context around user consumption patterns.
“I suspect the frontier AI models if I has a crystal ball they will spend a lot more time the next year or two building memory.”Nikesh Arora · 22 Jun 2026