AI-generated scientific discovery rates are substantially overstated; rigorous code review reduces confirmed novel findings from ~75% to ~30%.
The case
Web data is primarily attitudinal, not behavioral, and therefore insufficient for modeling human behavior.
“These models were trained on the web data like whatever was available in the web. And these are really interesting data sets, but they are fundamentally attitudinal data with some behavioral data that sprinkle around here and there.”Joon Sung Park · 21 Aug 2026
Self-consistency as a design metric is easily gamed when all generated proteins look identical, making high self-consistency scores untrustworthy on their own.
“It's easy to get self-consistency, consistent design of structures if all of your proteins look identical.”Matt McPartlon · 11 Aug 2026
Foundation models trained on descriptive data do not yet outperform linear models on causal, perturbational counterfactual tasks.
“Both us and many others in the field have found that these models that are trained on descriptive data do not yet outperform linear models on causal tasks, perturbational tasks, what we call counterfactual tasks.”Ci Chu · 21 Jul 2026
Meter's blog post found that about 50% of SWE-bench code that passes the benchmark test is completely unmergeable.
“Meter had this very interesting blog post where they were like about 50% of Sweepbench code that passes the Sweetbench test is completely unmergable.”swyx · 27 Jun 2026
AI cannot currently take a new hypothesis all the way to a scaled, finished material in consumer products.
“What I don't think you can do is go from like I have this new hypothesis to oh my gosh I have a new material it's scaled it's done it's in products in your iPhone you can't do that today.”Joseph Krause · 17 Jun 2026
Claude's leaked source code, written by Claude itself, was criticized by professional software architects and engineers as brittle.
“The code leaves and then if you looked at the code anybody who looked at the code who's a real software architect and engineer threw up. They were like, 'It made They're like, 'This stuff is brittle.'”Tony Fadell · 7 Jun 2026
The pushback
X-Cell accurately predicted perturbation effects in unseen active T-cells, including known TCR complex biology and putative T-cell inactivators found in screens, without ever training on that cell state.
“XL has not seen active cell T- cells and it's able to make accurate prediction not only on the known biology the TCR complex predicting their effect accurately that these are going to what we would expect to see but also it predicted the puditive T- cell inactivators that we found in the screen correctly as well.”Bo Wang · 21 Jul 2026
Auto-formalization using Lean caught and patched an implicit assumption in Robert Aumann's 50-year-old 'agree to disagree' theorem that had never been made explicit.
“Econ 101 there's this famous theorem agree to disagree by Nobel Prize winner Robert Olman and that is a 50-year-old theorem since 1976 everyone's been teaching it for 50 years there's an implicit assumption that was never made explicit that a prover was able to catch in the auto formalization process and was also able to patch the proof.”Karina Hong · 21 Jun 2026
The first version of the ESM Atlas was used by Fang's group to discover a new gene editing system.
“Actually, the first version of the ESM atlas was used by Funang's group to find a new gene editing system.”Alex Rives · 27 May 2026
Current AI models can produce scientific papers indistinguishable in quality from human-written papers.
“I think we now have models that can really turn out papers that are as good as human written papers.”Alex Lupsasca · 5 May 2026
AI could cure the majority of human diseases within the next decade (by ~2036).
“The prospect that we might cure the majority of human diseases in just the next decade or so is obviously extremely exciting.”Nathan Labenz · 1 Apr 2026