28 Aug 2026
Signal Headquarters
Vol. I
No. 262
· · 3 min read

AI benchmarks are saturating faster than the field can replace them

The canonical evaluations used to measure AI progress are worn out, and every attempt to build replacements is compromised almost the moment it ships. The gap between what models can do and what the field can measure is widening faster than most institutions have acknowledged.

Mark Chen puts the situation plainly: the field is in an “evals crisis.” The canonical benchmarks that the research community grew up using, including the SAT, are fully saturated. Worse, Chen adds that once an evaluation is released publicly, it is already no longer a good one. The two problems compound each other: the benchmarks that exist are worn out, and every attempt to replace them is compromised the moment it ships.

The numbers behind the saturation are stark. Brendan Foody notes that on the Apex benchmark, the frontier model now sits at about 40 percent. Twelve months ago, the frontier model scored 1 percent. Swyx projects that the Frontier Code 2026 benchmark will itself hit roughly 80 percent saturation by year’s end. The pace at which evaluations go from meaningful signal to ceiling effect has compressed from years to months.

The challenge of constructing tasks that are long enough to matter compounds the saturation problem. Nathan Labenz reports that evaluators at METR are “really struggling to have tasks long enough to even be able to evaluate these things.” Data Labenz cites shows a doubling time for AI task length of under four months, which implies an 8 to 12 times increase over the course of a single year. The overhead of designing evaluations at that scale is not something every benchmark team can absorb.

The proper way to evaluate the models now is you either have some kind of budget for the benchmark whether it's tokens or cost or time or whatever, or you plot the performance as a function of the amount of test time compute that's going into the model. Noam Brown

Some benchmarks are holding, but for reasons that cut the other way. Sergiy Nesterenko and Axel Backlund, discussing the Vending Bench simulation, note that models have improved enough to survive the full simulated year where earlier models crashed out. Yet Nesterenko estimates that a genuinely strong human performance would score roughly 10 times what current models achieve. Vending Bench is not saturated, but its headroom reflects the difficulty of the task, not confidence that the gap will persist.

The deeper structural problem is one Noam Brown identifies: a benchmark score is meaningless without specifying how much compute went into producing it. Current models, Brown notes, can think for weeks when scaffolded before performance plateaus. That makes any single score a function of spending as much as of capability. Brown is direct about what this requires: either evaluators fix a compute budget for the benchmark, or they plot performance as a function of test-time compute spent. Anything else produces numbers that cannot be compared across models. Brown also flags that inflating scores by scaffolding multiple model runs together is straightforward, which means that published numbers are not always what they appear. Brown makes the full implication explicit: if the goal is to know what a model can accomplish after running for a month, the only way to be sure is to actually run it for a month.

The safety implications have received less attention than the benchmark horse race, and Brown raises them directly. Preparedness frameworks and responsible scaling policies, he observes, do not account for test-time compute. They ask what a model can do, but in the current environment, capability is a function of how much money goes into inference. A framework built on static capability thresholds is measuring something that no longer maps cleanly onto risk.

Chamath Palihapitiya adds a dimension that makes the measurement problem harder to fix through institutional effort alone. Model performance advantages evaporate within weeks of publication, as open and closed competitors match or exceed frontier benchmarks almost immediately. Martin Casado puts the window at three to nine months of relevance for any given model. If evaluation infrastructure takes longer to build than a model’s competitive lifespan, the field is permanently behind. The benchmark ecosystem was designed for a slower world, and the clearest signal it is sending right now is that it cannot keep up.

The Editor, for the readers of Signal Headquarters

From the Archive