25 Aug 2026
Signal Headquarters
Vol. I
No. 237

AI models are outpacing available benchmarks, which can no longer construct tasks long enough to evaluate them.

The case

Anthropic's internal model 2 scores 8.8 percentage points higher on CobBench than Methos Preview, indicating a large gap between internal and publicly released model capabilities.

“On cobbench enthropic enthropics model 2 is about 8.8 percentage points higher than methos preview.”
Nathan Labenz · 22 Aug 2026

Poolside uses an internal process called 'polishing' to adapt models so they perform well on external harnesses, not just their own.

“We internally have been kind of this calling this polishing which is like you've got your model you do a little bit of polishing so that like it's able to work well on other harnesses as it is in your own.”
Eiso Kant · 22 Jul 2026

Frontier Code 2026 benchmark will be saturated at approximately 80% pass rate by end of 2026.

“Frontier Code 2026 will be saturated by the end of this year. We, you know, my estimate that you, we'll probably hit like 80% by the end of this year.”
swyx · 27 Jun 2026

Benchmark scores can be artificially inflated by scaffolding multiple model runs together, making cross-model comparisons misleading.

“It's really easy to show you can do much better than previous benchmarks or previous models on benchmarks by just for example scaffolding a bunch of models together.”
Noam Brown · 26 Jun 2026

The AI field is in an 'evals crisis' because the number of canonical gold-standard benchmarks is low and existing ones are saturated.

“We really are kind of in an evals crisis, right? Where all the really great EVAs that we all know like growing up like taking the SAT or those are all fully century and we really need to find good new ways to benchmark the models.”
Mark Chen · 25 Jun 2026

Radical AI's AI scientist has moved into elemental and alloy families that no one has ever published on before, demonstrating bias-free exploration beyond human scientists.

“Where our AI scientist has gone >> and it's moved into elemental families or alloy families no one has ever published on before.”
Joseph Krause · 17 Jun 2026

The pushback

State-of-the-art AI models still fail approximately 20% of the time on fourth-grade science tasks such as boiling water.

“The best models right now are getting something like 80% on the fourth grade science.”
Nathan Labenz · 6 Jun 2026

Current publicly accessible closed models are not good at selecting which experiment to run next in an automated research loop.

“What I find is that the current closed models the public can access today don't seem to be that great at selecting what the next experiment should be in a given track.”
Eric Jang · 15 May 2026

Internal documents show AI companies select which model capabilities to advance based on which industries will pay the most, choosing finance, law, medicine, and commerce rather than pursuing general intelligence.

“They create this myth that they are actually pushing the frontier of all of the capabilities of the model but that's not what's actually happening internally and I have I had hundreds of pages of documents on like how they were specifically training models they pick what capabilities they want to advance and you know how they pick them it's based on which industries countries would be able to pay them the most money for their services. So they pick finance, law, medicine, healthcare, commerce.”
Karen Hao · 26 Mar 2026

Topics

AI BenchmarksAI EvaluationFrontier AI

Signal Headquarters · compiled from attributed public discussion. Last updated 2026-08-22.