AI capability is compounding faster than institutions can track, and the commercial numbers prove it
Benchmark scores doubling in months, revenue growing 3x in a quarter, and development cycles measured in weeks: the pace of AI advancement is no longer a forecast. It is an operating condition.
The benchmark numbers are where to start, because they are the hardest to dismiss. Brendan Foody notes that frontier model scores on the Apex benchmark rose from 1% to 40% in 12 months. Nathan Labenz, analyzing the METR task-length data, puts the doubling time for AI task capability at under four months, which implies an 8 to 12x improvement over a single year. These are not measurements from the same benchmark or the same methodology. They point in the same direction.
On the model development side, Dylan Patel describes Anthropic’s capabilities advancing from L4 to L6 engineer level in two months. Patel also calls Anthropic’s Mythos model “potentially the biggest step up in model capabilities in like 2 years.” That framing is notable precisely because it comes from someone who tracks these systems closely: if Mythos registers as the biggest leap in two years to someone with Patel’s vantage point, the underlying pace of movement was already extraordinary before Mythos arrived. Eiso Kant offers a parallel data point on development tempo: Poolside completed a full model from the start of pre-training to launch in five weeks, then began pre-training on the next model the following day, illustrating that the build cycle itself is compressing at the same rate as the capability curve.
The security finding from Mythos is among the more concrete data points in the record. Krishna Rao reports that Anthropic’s Mythos found 250 security vulnerabilities in an open-source codebase where a prior model had found only 22. That is not a marginal recall improvement. It suggests the ceiling on what an automated audit can catch is moving faster than most security teams have accounted for.
From what we see, the scaling laws are not slowing down. Krishna Rao
Commercial velocity is running at the same tempo. Rao reports Anthropic’s annualized run-rate revenue grew from roughly $9 billion at the start of the year to north of $30 billion by the end of the most recent quarter. Amjad Masad describes Replit scaling from $2.5 million to $250 million in annual revenue in one year, with a path to $1 billion in the current year. Masad also notes that Replit’s daily ARR jumped from $1 million to $2 million in two days after launching Replit Agent. Marc Andreessen asserts that Anthropic and OpenAI are adding more revenue per month than Meta, Google, or Microsoft. Elad Gil adds that Anthropic and OpenAI each reached those heights in roughly a year. Chris Degnan captures the market’s new standard crisply: doubling revenue year-over-year, which would have been celebrated five years ago, may no longer be sufficient for survival.
The operational tempo that produces these numbers is itself striking. Laura Burkhauser says Descript will have a new model evaluated and integrated into its product within 15 minutes of release. Olive Song reports MiniMax ships a new model version approximately every month to a month and a half. Martin Casado argues that a given model remains relevant for three to six months, or six to nine months, before being superseded. Whatever the precise window, the implication is the same: enterprise software built around a stable foundation has no such foundation anymore. Gavriel Cohen makes the operational consequence explicit, noting that agents require constant model upgrades and that running a model version for years, as traditional enterprise software allows, simply does not work when the core system is changing continuously.
Beneath all of this sits a foundational question about whether the improvement rate can continue. Mark Chen argues that scaling laws have held for nearly 10 orders of magnitude with no sign of stopping. Rao states directly that, from what Anthropic observes, the scaling laws are not slowing down. Nathan Labenz adds that pre-training has continued to improve and was never actually plateauing in the way some commentary suggested. The debate on when the sigmoid bends is live, and no one inside the building is claiming to have a precise answer. What they are claiming, consistently, is that nothing in their current measurements shows a deceleration that registers in their revenue or their benchmark results.
What the evidence collectively describes is a situation where the traditional lag between capability advance and institutional adaptation is widening. The models are moving on month timescales. The businesses built on them are scaling at rates that have no historical SaaS analog. The benchmarks used to track progress are saturating before the measurement community can replace them. Dario Amodei puts a 50% probability on superintelligence by 2029. Whether or not that timeline proves accurate, the nearer-term data points suggest that the gap between what these systems can do and what most organizations have built to accommodate them is the more pressing problem.