Medical AI aces the textbook diagnosis and fails the messy real world start, new studies show

Started by Brad, Today at 12:42 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Medical AI aces the textbook diagnosis and fails the messy real world start, new studies show   Views(Read 74 times)
Active members in this topic:
Brad(1)

Brad

A cluster of recent studies is converging on an uncomfortable pattern for anyone hoping AI is about to transform frontline medicine any time soon, and the Financial Times has pulled the thread together under a fittingly blunt headline, medical AI has a proof problem. The technology's advances simply have not translated into meaningful improvements in real life care yet, and the specific reason why turns out to be more precise and more interesting than a generic complaint about AI hype outrunning substance.

A JAMA Network Open study published in April tested 21 different AI models against 29 standardised clinical cases and found something genuinely counterintuitive. Failure rates for differential diagnosis exceeded 80 percent when the models were given incomplete patient data, the messy, partial information state doctors actually work with early in a real consultation, but that failure rate dropped below 40 percent once the models were fed complete information instead. Arya Rao, the study's lead author at Mass General Brigham, summed up the trap neatly, noting that these models are great at naming a final diagnosis once the data is complete, but they struggle at the open ended start, which is precisely the part of medicine that actually requires skilled judgement under uncertainty.

A separate study published in PLOS Digital Health exposes an even more structural problem sitting underneath the performance numbers. Of 1,357 AI medical devices that have received FDA clearance, only 3 of them, a mere 0.2 percent, have ever undergone patient centred outcome testing looking at things like mortality or hospital readmission rates. Radiology, one of the fields most aggressively adopting AI tools, fares no better, with just 3 out of 1,059 cleared devices registering prospective clinical trials, itself only 0.3 percent of the total. On top of that thin evidence base, 62 percent of the existing research relied on small, homogenous patient cohorts that routinely excluded pregnant women, children and non English speakers, groups whose absence from the validation data makes any confident claims about general real world performance considerably shakier than the marketing around these tools tends to suggest.

The regulatory gap driving all of this gets spelled out clearly by Sebastián A Cajas Ordóñez from MIT, who points out that FDA clearance means a device is substantially equivalent to something already on the market, not that using it actually makes patients better off. That distinction, between regulatory equivalence and genuine clinical benefit, is exactly the gap that lets a medical AI tool reach the market and get deployed widely without ever having to prove it improves outcomes for the people it is actually treating.

What makes this particular critique land harder than the usual AI skepticism is how broad the sample was. The studies tested systems from OpenAI, Google, Anthropic, xAI and DeepSeek, meaning this is not a story about one underperforming vendor falling behind its competitors, it is a structural problem running across essentially every major AI lab's medical applications simultaneously, which points squarely at how these systems get validated and approved rather than at any particular company's engineering shortcomings.
Tapped out by my own semicolon again

Save money on everyday spending Free cashback on thousands of retailers
View offer