Custom fine-tuned models deliver significantly lower latency and cost than frontier models while maintaining equal or higher quality.
The case
A 10-to-1 price discrimination ratio from frontier model providers makes it effectively impossible for app-layer companies to compete while using those providers' APIs.
“At a 10 to1 price discrimination ratio it's just really hard for them to compete.”Nathan Labenz · 22 Aug 2026
Some capabilities in human behavior modeling cannot be achieved through prompting alone; model parameters must be modified.
“There are certain things you just cannot shape just by prompting the model. So some to some degree you do need to touch the parameters of the model itself.”Joon Sung Park · 21 Aug 2026
Grok 4.5 uses just one-third the tokens of GPT-5.5 or Fable while achieving a similar intelligence score.
“Grok 4.5 uses just one-third the amount of tokens as GPT-5.5 or Fable while achieving a similar score.”Ryan Greenblatt · 11 Aug 2026
Silico's $1,000/month price will decrease over time as Goodfire develops more token-efficient agent methods.
“My hope is we're starting out with $1,000 a month subscription. We'll be able to bring that down over time because we're able to come up with more and more clever ways.”Dan Balsam · 8 Aug 2026
Open-weight models can reach speeds of up to ~500 tokens per second, which is 2-3x faster than the proprietary fast mode available today.
“For proprietary model there is regular mode and fast mode and that's only the two switch here. But for openweight when you're running it, every provider can offer potentially even 10 different levels of speed going from like the slowest mode which can be a lot cheaper to 400 tokens per second almost up to 500 in many cases that for some workloads and this is typically 2x or 3x faster than the fast mode out there today.”Simon Mo · 6 Aug 2026
Fine-tuned smaller models can be simultaneously better, cheaper, and faster than frontier models on specific tasks.
“We end up getting all three things. It is better at the toss. It is cheaper and it is faster.”Jesse Zhang · 31 Jul 2026
The pushback
Anthropic has become a premium product since the end of last year, priced at twice the cost of its competitor.
“It has become a premium product since the end of last year. It is twice the cost of its competitor, right?”Jason Lemkin · 28 May 2026
A vulnerability discovered approximately two weeks prior in LiteLLM was found to steal all user keys and credentials, making it unsafe for enterprise deployment.
“We have seen with light LLM, how long is this now ago, 2 weeks or something? You probably saw it, right? Like with this vulnerability that still all of a sudden steals all your keys and credentials and so on and so forth.”Philipp Herzig · 23 Apr 2026