AI has already reached specialist-level performance, and institutions have not caught up
Geoffrey Hinton points to a Microsoft blog describing AI instances that outperform most doctors in diagnosis. That claim is secondhand. What surrounds it is not: experimental troubleshooting at PhD level, research compression from 160 years to a $10,000 budget, and a credible projection of expert-surpassing AI by the early 2030s.
Geoffrey Hinton, the Nobel Prize-winning AI researcher, points to a Microsoft blog post describing a result that would have seemed implausible not long ago: multiple AI instances, interacting with each other, outperform most doctors in medical diagnosis. Hinton is citing a blog rather than a peer-reviewed study, and that caveat matters. The claim should be read as a reported finding rather than an established benchmark. But the direction he is pointing is not his alone.
What makes it harder to dismiss is how many other domains are producing similar findings at the same time.
Geoffrey Irving, a researcher focused on AI safety, describes what frontier models can now do in scientific settings: deliver troubleshooting advice at a level described as PhD-quality, working from nothing more than a photograph of an experimental setup and a small amount of accompanying text. That capability, applied at scale, would change the economics of experimental science in the same way that Hinton’s medical example changes the economics of specialist consultation. Ryan Greenblatt puts the machine learning research domain in similar terms, noting that AI systems can already match humans who are mediocre at machine learning research at doing that work. The floor of what AI can do is no longer amateur-tier.
David Sinclair, whose laboratory focuses on the biology of aging, offers the most concrete compression figure in the evidence. Work that would previously have required 160 years and, in his words, “quite literally billions of dollars” is now being completed on a $10,000 budget. That is not a productivity gain that fits comfortably inside existing institutional structures. It is a change in what a small team can attempt at all.
The prospect that we might cure the majority of human diseases in just the next decade or so is obviously extremely exciting. Nathan Labenz
The behavioral shift is already visible at the working level. Tulsee Doshi, a researcher at Google DeepMind, described a colleague running complex ablation experiments on Gemini models from her phone, in a hot tub, over the course of an hour, producing a full report at the end of it. The anecdote is casual, which is precisely what makes it informative: the capability has become routine enough that it happens incidentally rather than as a deliberate demonstration.
The anecdotal and the quantified point the same direction on trajectory. Ajeya Cotra, a researcher who works on AI forecasting, projects that the early 2030s will see what Ryan Greenblatt calls “top human expert dominating AI,” meaning systems that surpass any human expert at remote tasks. That framing is forward-looking, and projections about AI capability timelines have a poor historical track record in both directions. But Cotra’s projection does not require a discontinuous leap from where the evidence already sits. It requires the current rate of improvement to continue for roughly half a decade.
Sam Parr, the entrepreneur and media founder, has referenced a person from GitLab who, by his account, used AI to cure his own cancer. That claim arrives secondhand and its clinical specifics are not supplied. Taken alongside Hinton’s diagnostic example, it fits a pattern of people beginning to use AI as a first-order medical resource rather than a supplementary one. Whether that is wise depends heavily on the specific case. That it is happening is the relevant data point here.
Nathan Labenz frames the endpoint of this trajectory directly, calling the prospect of curing the majority of human diseases within the next decade “obviously extremely exciting.” No current evidence can confirm that outcome. What the evidence can support is the narrower claim: AI has already performed at or above specialist level in diagnosis, laboratory troubleshooting, and domain-specific research tasks, and the institutions built around the assumption that such work requires years of human training and large teams have not yet reckoned with what that means. The force multiplier is operating now. The institutional response has not caught up.