11 Aug 2026
Signal Headquarters
Vol. I
No. 188
· · 4 min read

AI capability is compounding faster than the institutions built to measure it can track

Benchmark scores are doubling in months, revenue is tripling in quarters, and the people building these systems are now shortening their own timelines. The data points, taken together, describe something beyond an acceleration: they describe a structural change in what is possible and how fast.

The benchmark numbers are where the case starts. Brendan Foody notes that frontier model scores on the Apex benchmark rose from 1% to 40% in 12 months. Nathan Labenz, citing a METR chart, puts the doubling time for AI task length at under four months, which implies an 8 to 12x annual gain. The Stanford 2026 AI Index corroborates the direction from outside: frontier models gained 30 percentage points in a single year on Humanity’s Last Exam, a benchmark designed to be hard for AI and favorable to human experts. Evaluations built to hold for years are saturating in months. swyx estimates the Frontier Code 2026 benchmark will be saturated at roughly 80% by the end of this year.

The commercial numbers are just as stark. Krishna Rao reports that Anthropic ended its most recent quarter at north of $30 billion in annualized run-rate revenue, up from roughly $9 billion at the start of the year. Marc Andreessen observes that Anthropic and OpenAI are adding more revenue per month than Meta, Google, or Microsoft, and projects a combined run rate of $200 billion for the two companies by the end of 2026. Amjad Masad reports that Replit grew from $2.5 million to $250 million in annual revenue in a single year, tracking toward $1 billion in the current year, and that on the day Replit Agent launched the product added $1 million in ARR in its first 24 hours and $2 million in its second. Elad Gil notes that both Anthropic and OpenAI reached $1 billion in revenue in approximately a year. Dylan Patel reports that Anthropic’s models advanced from L4 to L6 engineer level in two months.

Specific task performance tells the same story in concrete terms. Krishna Rao describes Anthropic’s Mythos model finding 250 security vulnerabilities in an open-source codebase where a prior model had found only 22. Carina Hong reports that Axiom Math scored a perfect 120 out of 120 on the 2025 Putnam exam, surpassing DeepSeek’s best LLM score of 103 and the best human score of 110. Hong also notes that Axiom’s underlying prover has scaled from handling proof trees of 40 nodes to 4,000 nodes. Cat Wu observes that multi-agent simultaneous code review across an entire codebase only became reliable enough for production use with recent model versions, and that token cost per engineer rises with each model jump as workers delegate more tasks. These are not marginal improvements in existing workflows. They are capability thresholds crossed.

We started the year with about $9 billion of run rate revenue and we ended the quarter with, you know, north of $30 billion of run rate revenue. Krishna Rao

The scaling laws underpinning all of this are, by the accounts of the people closest to them, intact. Mark Chen reports that scaling laws have held for almost ten orders of magnitude and sees no reason they should stop holding. Krishna Rao says that from what Anthropic observes, scaling laws are not slowing down, contradicting a widely held belief that they are hitting limits. Nathan Labenz adds that pre-training never stopped working. Patel describes Mythos as potentially the biggest step up in model capabilities in roughly two years. David Dalrymple notes that each new model, with a couple of recent exceptions he names explicitly, is moving in the direction of being not just more intelligent but more capable of sound judgment.

The velocity of the development cycle is itself a signal. Poolside’s Eiso Kant reports that the Laguna XS2 model went from the beginning of pre-training to launch in five weeks, and that the next model began pre-training the day after post-training on the previous one concluded. Martin Casado observes that a given model remains relevant for only three to nine months before being superseded. Gavriel Cohen notes that agents, unlike traditional enterprise software that can sit unchanged on a server for years, require constant model upgrades because the core thing being built on is constantly changing. Laura Burkhauser describes Descript as able to evaluate and integrate a new model from a major lab within 15 minutes of release.

The people making these systems are shortening their own timelines. Dario Amodei puts the median probability of superintelligence at 50% by 2029, and reports that people inside AI companies are now pushing that date back to 2027 or 2028. Mo Gawdat expects AGI, meaning AI capable of performing most human tasks better than humans, by end of 2027 at the latest. Sarah Guo characterizes coding as essentially a solved problem by the end of this year and expects a light form of recursive self-improvement by end of 2027. Sebastian Mallaby projects recursive self-improvement, where the frontier model codes the next frontier model itself, by 2028. These are not fringe forecasts from outside observers. They come from investors, researchers, and company founders with direct access to what current systems can do.

What emerges from the full set of evidence is not just a speed story. Replit’s Masad notes that OpenAI and Anthropic researchers told him they did not know their own models were capable of end-to-end coding as demonstrated by Replit Agent. If the developers of the systems are surprised by what the systems can do, the question of where the ceiling sits is genuinely open. The data points on benchmarks, revenue, task performance, and timelines do not contradict each other. They point in the same direction, and the direction is not slowing down.

The Editor, for the readers of Signal Headquarters

AI BenchmarksAI Capability GapAI Model ReleasesAI PerformanceAI Revenue GrowthAI Timelines



From the Archive