Most FDA-cleared AI medical devices, including those used in radiology, have not been tested to demonstrate they actually improve patient health outcomes. Out of over 1,300 FDA-authorized AI medical devices, only three have been evaluated for their impact on patient outcomes like mortality, morbidity, or hospital readmissions.
- Only 3 out of 1,357 FDA-cleared AI devices (0.2%) were tested on patient-centered outcomes like mortality or readmissions
- In radiology specifically, 3 out of 1,059 cleared AI devices had registered prospective trials, representing just 0.3% of tools
- FDA clearance through the 510(k) pathway means a device is substantially equivalent to existing products, not that it improves patient outcomes
- 62% of AI device studies used small, homogenous groups and frequently excluded vulnerable populations like pregnant women, children, and non-English speakers
- Experts recommend a three-phase evidence ladder requiring prospective studies of at least 500 patients before clearance and multi-center trials of 2,000+ patients after clearance
Most AI medical tools, including those used in radiology, were not tested on patient outcomes prior to regulatory clearance, according to a report published August 19 in PLOS Digital Health.
Out of the 1,300-plus AI-based medical devices authorized by the U.S. Food and Drug Administration (FDA) for use in patient care, only three had been tested on whether they actually improve patients’ health, wrote a team led by the Massachusetts Institute of Technology (MIT) Critical Care team
“We expected the evidence base to be thin. We did not expect three,” said co-author Sebastián A. Cajas Ordóñez to AuntMinnie.
AI continues to advance into clinical practice, with 1,357 AI and machine learning-based medical devices being either cleared or approved by the FDA. Over three out of four FDA-cleared AI devices are used in radiology. Ordóñez said radiology has the thinnest prospective evidence base of any major specialty, with registered trials for 0.3% of its cleared tools.
The MIT Critical Care team analyzed these devices, approved or cleared through December 5, 2025, using the FDA device database and the American College of Radiology (ACR) Data Science Institute catalogue. The team identified registered trials and publications via linked searches of ClinicalTrials.gov and PubMed.
Of the total cleared AI devices, 34 (2.5%) were linked to registered prospective trials. Twelve devices (0.9%) posted results and 12 (0.9%) had peer-reviewed publications.
Only three devices (0.2%) evaluated patient-centered outcomes such as mortality, morbidity, or readmissions. Of the 1,059 cleared devices used in radiology, three had a registered trial.
Finally, 62% of studies used observational designs with “small, homogenous cohorts, limited subgroup analyses, and frequent exclusion of vulnerable populations.”
The study authors suggested that structural barriers may hinder rigorous evaluation. These include misaligned financial incentives, reliance on predicate-based regulatory pathways, and logistical challenges of multi-center trials.
“Under the 510(k) pathway, it [clearance] means a device is substantially equivalent to something already on the market, not that using it makes patients better off. Clearance is the beginning of due diligence, not the end,” Ordóñez told AuntMinnie. “The full device list and trial identifiers are published with the paper, so a radiologist can look up the tool in their own reading room and see exactly what public evidence stands behind it.”
Ordóñez said before buying these AI devices, practices should ask vendors three things: what was the trial registration number and primary endpoint, does the validation population resemble ours, and who was excluded?
“Pregnancy was an exclusion in a third of radiology trials and 42% of cardiovascular ones, pediatric patients were excluded almost universally, and non-English speakers were frequently omitted,” he said. “Those patients still show up in your emergency department.”
The MIT Critical Data team suggested a three-phase stage evidence ladder could address these issues.
“Phase 0, before clearance, retrospective validation on diverse datasets with mandatory demographic reporting; Phase 1, around clearance, prospective studies of at least 500 patients embedded in real workflows; Phase 2, after clearance, multi-center outcome trials of at least 2,000 patients with pre-specified subgroup analyses,” it suggested. “The 500-patient bar is where 73.5% of today's trials already fall short.”
However, it added that the absence of public evidence “is not proof the FDA is asleep.”
“The agency has non-public mechanisms such as manufacturer reporting and MAUDE [Manufacturer and User Facility Device Experience], but clinicians cannot read those, and that transparency gap is a separate, real problem,” the team said.
“And the FDA is not the only lever. Journals can enforce Spirit-AI and Consort-AI, and payers can tie reimbursement to demonstrate benefit rather than regulatory status,” the team added. “The single most immediate change at the agency would require pre-registration of prospective trials for Class II and III AI devices. None of this is a call to slow innovation. It is a call to make it count.”
Read the study here.



















