Less than 30% of FDA-authorized radiology AI devices have undergone clinical testing
New data further highlight ongoing concerns about a lack of prospective clinical testing for artificial intelligence and machine learning-enabled applications authorized by the U.S. Food and Drug Administration.
The FDA has approved more than 1,000 AI/ML devices, the majority of which are radiology-specific. As such, many experts have expressed concern about how such rapid growth could affect patient safety and outcomes, as the FDA does not have predefined standards for efficacy, safety and risk reporting of AI/ML devices prior to or after clearance.
Published in JAMA Network Open, a new paper details the path to approval for nearly 1,000 AI/ML algorithms, highlighting significant gaps in testing. For example, of the 723 devices that had been approved for radiology applications at the time of the study, less than 30% underwent clinical testing, and even fewer were subject to prospective testing.
The authors of the paper caution that, as AI/ML devices further infiltrate clinical settings, stringent oversight will become critical.
“Artificial intelligence and machine learning increasingly support clinical decision-making, particularly in radiology. The number of AI-enabled tools cleared by the US Food and Drug Administration continues to rise,” Ram Sivakumar, MD, of Washington University in St. Louis, and colleagues noted. “However, evidence about clinical generalizability is lacking. We demonstrate that device testing gaps underscore the need for clinical oversight.”
For their analysis, the group reviewed all AI/ML-enabled devices that earned FDA pre-market authorization from November 1995 to June 2024. They compared information pertaining to device evaluation, such as prospective, human-in-the-loop and clinical testing, paying close attention to testing methods for devices specific to radiology.
At the time of the study, there were 950 FDA-authorized AI/ML devices, 723 of which targeted radiology. The majority of these devices (97%) were cleared via the 510(k) pathway, which streamlines market entry; this requires developers to ensure that their product is equivalent to a legally marketed device and is equally safe and effective. However, this pathway does not require clinical data demonstrating performance or safety, the authors noted. Another 22 were authorized as de novo applications (no predicate device required), while 4 received pre-market approvals, indicating higher associated risks.
In terms of testing, just 5% of the devices underwent evaluations that were prospective in nature, while 8% included a human-in-the-loop and 29% incorporated clinical testing. Just 15 devices included prospective and clinical testing, and 6 utilized all 3 methods.
The team suggested that the lack of observed involvement from human operators during testing could limit devices’ clinical utility.
“Today, most AI/ML devices are used in conjunction with a human, yet only 56 were tested with any human operator. Most have not been validated against defined clinical or performance end points,” the group noted. “In a study of radiologists interpreting chest radiographs with AI assistance, high performers maintained strong performance, while low performers did not necessarily improve. Such heterogeneity in performance raises questions about clinical generalizability.”
Current regulations put more emphasis on safety over clinical utility, but the authors contend that utility is just as important. Clinical studies and AI frameworks similar to those in Europe (Health Technology Assessment) could help address evidence gaps relative to utility, the group suggested.
Read more here.
