Could ChatGPT be up to the task of monitoring AI drift?

Artificial intelligence tools in radiology require constant monitoring to ensure model performance remains consistent. New research suggests OpenAI’s ChatGPT could be up to the task. 

Implementing AI into clinical practice is only the beginning of the AI journey, as algorithms require oversight to monitor for performance glitches. One of the biggest concerns regarding the use of AI in medicine is the degradation of performance over time. "AI drift” occurs due to a variety of factors; catching variations in its outputs before they lead to inaccurate or misleading predictions is critical to ensure model reliability and prevent patient harm. 

“Traditional drift detection approaches, which rely on real-time feedback, are often impractical in healthcare settings due to delays in obtaining ground-truth data,” Mohammad Ghasemi-Rad, MD, with the department of radiology at Baylor College of Medicine, and colleagues explained in Academic Radiology. “This limitation creates a significant challenge for accurate and timely performance monitoring, potentially compromising the safe and effective use of AI in clinical workflows.” 

Keeping tabs on AI performance, though necessary, is a challenging task. Given that many organizations are already struggling to address their current workloads, adding more responsibilities to staff may not be feasible. Large language models like ChatGPT may represent a solution to the issue by analyzing radiology reports for patterns indicative of model drift. 

Subscribe to Radiology Business News

The group recently tested a HIPAA-compliant Microsoft Azure-hosted version of ChatGPT-4 Turbo on a set of non-contrast head CT radiology reports generated by an AI tool from Aidoc. The imaging exams were conducted to rule out intracranial hemorrhage (ICH), and GPT-4 Tubro was tasked with extracting data from the reports to compare against a ground truth and identify patterns in performance. 

GPT had achieved an accuracy of 99.5% for identifying report alterations. When compared with the outputs of the Aidoc system, the LLM’s extracted ICH information yielded a 60% concordance rate.  

Of the reports flagged by GPT for containing false positives, researchers were able to identify several factors that had contributed to the AI system’s mistake. The misclassifications were influenced by scanner manufacturer, midline shift, mass effect, artifacts and patients’ neurologic symptoms. 

“This study underscores the importance of continuous performance monitoring for AI systems in clinical practice,” the authors noted. “The integration of LLMs offers a scalable solution for evaluating AI performance, ensuring reliable deployment and enhancing diagnostic workflows.” 

Although more longitudinal research is needed before LLMs can be trusted with AI monitoring, the authors maintain that the method offers a cost-friendly and promising way to keep algorithms in check. 

Read more about the research here.

Hannah Murphy
Hannah Murphy, Editor

In addition to her background in journalism, Hannah also has patient-facing experience in clinical settings, having spent more than 12 years working as a registered rad tech. She began covering the medical imaging industry for Innovate Healthcare in 2021.

Subscribe to Radiology Business News

Subscribe to Radiology Business News