AI model capabilities are advancing extremely rapidly, with benchmark scores and revenue doubling in months.
The case
Anthropic's internal model 2 scores 8.8 percentage points higher on CobBench than Methos Preview, indicating a large gap between internal and publicly released model capabilities.
“On cobbench enthropic enthropics model 2 is about 8.8 percentage points higher than methos preview.”Nathan Labenz · 22 Aug 2026
The median 2026 Stripe founding cohort generates 50% more revenue than the comparable 2025 cohort.
“The median 26 cohort is generating 50% more revenue than the comparable 25 cohort.”Patrick Collison · 17 Aug 2026
Grok 4.5 uses just one-third the tokens of GPT-5.5 or Fable while achieving a similar intelligence score.
“Grok 4.5 uses just one-third the amount of tokens as GPT-5.5 or Fable while achieving a similar score.”Ryan Greenblatt · 11 Aug 2026
Frontier AI models are causing a massive reduction in the time between vulnerability discovery and exploitation.
“They are causing kind of a massive reduction in the time between the vulnerability discovery and vulnerability exploitation.”Zane Lackey · 7 Aug 2026
A light form of AGI (RSI) will be achieved by the end of 2027.
“We're 6 monthsish or towards the end of the year to be completely done with code. Like it's a solved problem and then we'll probably hit some form of, you know, light RSI by end of next year.”Sarah Guo · 6 Aug 2026
The pushback
GPT-4.5 was internally considered a bust by people at OpenAI.
“There's GPT-4.5, which famously people at OpenAI thought was a bit of a bust.”Ryan Greenblatt · 11 Aug 2026
Fable model showed a regression compared to Opus on subjective nuanced tasks, demonstrating that greater model power does not always improve performance in these domains.
“Fable was actually a big regression, interestingly. which I think is just an interesting point which shows that more power doesn't like when you when you're dealing with these sort of more subjective nuance things, more power doesn't always mean better.”Cameron · 27 Jun 2026
State-of-the-art AI models still fail approximately 20% of the time on fourth-grade science tasks such as boiling water.
“The best models right now are getting something like 80% on the fourth grade science.”Nathan Labenz · 6 Jun 2026
Coding models are approaching a performance plateau, making fine-tuning a viable strategy for use-case-specific optimization.
“We're approaching a certain plateau in how good coding your data to fine-tune a model specifically for your use case.”Amjad Masad · 25 Apr 2026
Reinforcement learning is not ready to learn PCB routing from scratch by interacting with CAD software via keyboard and mouse.
“The reality is that I don't think reinforcement learning as a technology is ready for something.”Sergiy Nesterenko · 15 Apr 2026
Internal documents show AI companies select which model capabilities to advance based on which industries will pay the most, choosing finance, law, medicine, and commerce rather than pursuing general intelligence.
“They create this myth that they are actually pushing the frontier of all of the capabilities of the model but that's not what's actually happening internally and I have I had hundreds of pages of documents on like how they were specifically training models they pick what capabilities they want to advance and you know how they pick them it's based on which industries countries would be able to pay them the most money for their services. So they pick finance, law, medicine, healthcare, commerce.”Karen Hao · 26 Mar 2026