LLMs could automate quality control processes, saving radiology departments significant time
New data support the role of large language models (LLMs) in streamlining quality control initiatives.
Published in the Journal of the American College of Radiology, the findings suggest that LLM-based systems can automate the processing of radiology report checks. The study specifically focused on identifying variability in breast ultrasound reports, but experts believe these systems have potential beyond breast imaging alone.
“Trained on extensive medical text corpora, LLMs have demonstrated proficiency in understanding clinical terminology, contextual relationships, and logical reasoning,” corresponding author Hongyan Wang, with the department of ultrasound at Peking Union Medical College Hospital in Beijing, China, and colleagues noted. “These strengths align closely with the requirements of report [quality control], enabling LLMs to automatically identify errors, inconsistencies, and omissions in US reports in accordance with established reporting standards such as BI-RADS.”
To determine how an LLM would handle quality control, the group retrospectively analyzed breast ultrasound reports from 735 patients with pathologically confirmed breast masses across 60 hospitals in China. Investigators compared the performance of the Qwen2.5-VL-7B LLM with that of hospital quality control personnel, who manually converted free text reports into standardized, BI-RADS-based structured reports. Accuracy and time were both used to determine performance.
The LLM outperformed manual reviewers in evaluating reports for several breast lesion characteristics, including margins (80% versus 64%) and echo patterns (74% versus 56%). Its performance persisted in more complex reports involving multiple lesions, and it improved as BI-RADS categories increased from 3 to 5, indicating particularly strong performance in reports containing more suspicious findings. What's more, the model delivered a substantial workflow benefit, completing quality for 50 reports in an average of 13 minutes compared to 213 minutes for manual reviewers.
“Importantly, even in the more complex context of multi-lesion reports, the LLM demonstrated accuracy comparable to manual [quality control], highlighting its potential for processing complex clinical narratives,” the authors noted. “This advantage likely stems from the LLM’s extensive pretraining on medical texts, strong contextual understanding, and resistance to fatigue, ensuring stable and consistent QC performance.”
Although the team believes LLMs are promising for quality control in the future, they acknowledge the use of such technology is accompanied by concerns related to data privacy and the need to be frequently updated with new information. They suggested future studies should evaluate user acceptance, infrastructure requirements and long‑term model stability in real world settings.
Read more here.
