25 Jul 2026
Signal Headquarters
Vol. I
No. 148
Signal
· · 3 min read

AI scientific productivity is substantially overstated once the code behind the discoveries is actually checked

The headline numbers on AI-generated scientific discovery look impressive until someone examines the underlying work. Rigorous code review has collapsed confirmed novel findings to roughly 30 percent of initial claims, and the problem runs across domains from materials science to automated research loops.

The headline numbers on AI scientific productivity are striking. The numbers that survive rigorous scrutiny are considerably less so. Nathan Labenz offers the starkest illustration available: after someone spent days combing through thousands of lines of code generated by AI models to support their claimed discoveries, the share of findings that held up dropped to roughly 30 percent. That is not a rounding error. It is a systematic gap between what these systems appear to produce and what they actually deliver once the work is checked.

The pattern shows up outside pure science as well. Swyx points to a finding from Meter’s engineering blog: about 50 percent of code that passes SWE-bench benchmark tests is completely unmergeable in practice. A benchmark pass, in other words, does not mean working software. It means software that cleared a specific test under conditions that may not reflect what production deployment actually requires. The gap between benchmark performance and real-world utility is the same structural problem that Labenz’s audit exposed in scientific research, just in a different domain.

The problem is not limited to code quality. Joseph Krause, working in materials science, draws a sharper boundary: the bottleneck in materials discovery is not AI candidate generation but everything that comes after it. Synthesis, characterization, and manufacturing are where new materials either prove out or fail, and no single model can take a hypothesis all the way to a scaled material ready for consumer products. As Krause puts it, that is simply not how materials work. The implication is that AI accelerates one stage of a multi-stage pipeline while the other stages remain governed by physical and engineering constraints that computation does not dissolve.

Then we convinced somebody me to go through and spend days and days and days looking at the thousands upon thousands of lines of code that these models were generating to support their discoveries. and it went down to 30% of the discoveries were probably real. Nathan Labenz

Ci Chu adds a more foundational limitation. Foundation models trained on descriptive data do not yet outperform linear models on causal, perturbational, or counterfactual tasks. That is a significant constraint for any scientific application where the goal is not to describe a pattern that already exists in training data but to predict what would happen under conditions the model has never seen. Descriptive competence and causal reasoning are different capabilities, and the gap between them matters most precisely in the research contexts where AI is being promoted most aggressively.

Eric Jang identifies the same limitation at the experimental design level. Current publicly accessible closed models are not reliable at selecting which experiment to run next within an automated research loop. That is a critical weakness for the vision of AI-driven autonomous science: a system that cannot reliably choose the right next step cannot close the loop between hypothesis and result without significant human intervention at each decision point.

Alex Lupsasca names what all of this points toward. The verification step, he argues, is going to become a larger bottleneck this year. That framing is important: the problem is not static. As AI systems generate more candidate discoveries faster, the human capacity to audit those claims at equivalent speed does not scale with them. The asymmetry between generation and verification is structural, and it compounds over time. Bradley Sutton’s observation about the app economy captures the same dynamic from a different angle: app releases rose by nearly 80 percent compared to the prior year, while apps with significant usage and reviews remained flat or declined. Shipping got easier. Building things people actually use did not.

None of this means AI is not useful in research settings. The more precise reading is that its usefulness is concentrated at specific stages of work and degrades sharply when the output is not subjected to the same scrutiny applied to human-generated research. The field is currently running on a significant gap between measured performance on narrow tasks and actual scientific contribution. Closing that gap requires treating verification as a first-class problem, not an afterthought to generation.

The Editor, for the readers of Signal Headquarters

From the Archive