Large language models outperform physicians at imaging modality selection, study shows

Large language models may sometimes be better than medical providers at choosing the appropriate modality for patients in need of imaging. 

New data published in Clinical Imaging detail the potential of four different LLMs for determining the best imaging pathway for patients based on 120 different clinical scenarios. Experts determined that not only could the LLMs match the performance of clinicians, in many cases, they could surpass them as well. 

The findings have renewed optimism around how LLMs can positively impact imaging utilization, temper resource waste and optimize radiation doses for patients, authors of the paper suggest. 

“Conventional rules-based, keyword-matching engines (Google, Bing, Yandex, etc.) and simple text parsing techniques that have existed for decades can extract predefined terms from requisition text, but they are inherently brittle: they fail when synonyms, abbreviations, negations, or competing different clinical factors that are critical for determining patient-based appropriate imaging modality appear, and they cannot weigh radiation, contrast risk, and purpose of current imaging in an integrated manner,” Eren Çamur, with the department of radiology, Ministry of Health Ankara 29 Mayis State Hospital in Turkey, and co-authors noted. “LLMs, by contrast, derive a semantic representation of the entire request, permitting contextual reasoning across multiple variables and dynamic alignment with the evidence hierarchy embedded in the American College of Radiology Appropriateness Criteria (ACR AC).” 

Subscribe to Radiology Business News

Researchers used three different prompts on four LLMs—DeepSeek-V3, ChatGPT-4o, Claude 3 Opus, Claude 3.5 Sonnet—to determine their reliability in making imaging recommendations for 120 “multifaceted practice-oriented clinical scenarios.” These scenarios covered a wide range of concerns, including breast, cardiac, gastrointestinal, musculoskeletal, neuro, thoracic, genitourinary and vascular. The LLMs’ responses were compared to that of four clinicians—emergency, cardiologist, internist and general surgeon—and four radiologists with varying levels of experience to judge performance. All responses were categorized according to ACR AC. 

For all three prompts, all LLMs provided the same modality recommendations. DeepSeek-R1 led the group in both ACR clinical scenarios, at an accuracy of 98.3%, and in realistic scenarios, for which it matched the performance of junior radiologists and exceeded that of the other clinicians and residents. That same model, which was introduced in December, yielded short- and long-term reproducibility ranging from 0.773 to 0.886 and 0.507 to 0.787, respectively. 

“This study represents the first comprehensive evaluation of DeepSeek in radiology, directly comparing multiple LLMs to radiologists with different experience levels under both standardized and multifactorial conditions,” the authors noted. “In guideline-driven ACR scenarios, each model paralleled the accuracy of board-certified junior radiologists. When confronted with MPCS marked by comorbidities, polypharmacy, atypical presentations, and evolving laboratory findings, DeepSeek-R1 significantly outperformed clinicians and residents, while still achieving parity with both junior and senior radiologists.” 

The group suggested their findings support the “remarkable potential” of LLMs to improve radiology workflows.  

Learn more about the LLMs’ performance here. 

Hannah Murphy
Hannah Murphy, Editor

In addition to her background in journalism, Hannah also has patient-facing experience in clinical settings, having spent more than 12 years working as a registered rad tech. She began covering the medical imaging industry for Innovate Healthcare in 2021.

Subscribe to Radiology Business News

Subscribe to Radiology Business News