Factors that make radiologists less likely to be fooled by large language models

New research explores some of the key factors that make radiologists less likely to be fooled by incorrect advice from large language models. 

LLMs such as Claude or ChatGPT have shown strong capabilities in solving diagnostic problems and explaining their logic. However, their outputs aren’t always reliable, with these AI tools sometimes producing factually incorrect explanations with undue confidence, omitting key findings or overgeneralizing patterns, experts write Tuesday in RSNA’s Radiology

South Korean researchers recently sought to suss out the determining factors that might lead to radiologists getting duped. Their retrospective study involved 10 readers interpreting chest imaging from 100 patients. They found that successful radiologist-large language model collaboration was associated with confidence in the AI model, along with the reader possessing expertise in chest imaging. 

“Although high model confidence was associated with correct decisions, expertise may serve as a safeguard against persuasive but incorrect rationales,” corresponding author Taehee Lee, MD, MSc, with the Department of Radiology at the Seoul National University Hospital and College of Medicine, and colleagues wrote Aug. 11. “The double-edged role of rationale quality suggests explainability benefits those able to critically appraise it. Overall, effective LLM assistance depends on factors beyond model performance, including reader expertise and model confidence.”

The study utilized curated data from the Korean Society of Thoracic Radiology Weekly Case platform, spanning 2018 to 2020, including X-rays, CTs, MRIs and PET scans. Lee and colleagues compared readers’ handling of one session without AI against a second session randomized to use 1 of 2 different LLM models. The latter included a high accuracy (76%) model deployed OpenAI’s GPT-5 or a low accuracy one (27%) employing GPT-4o. Readers were provided multiple-choice diagnostic options with rationales against a reference standard, established by the case author. 

Model confidence (odds ratio of 3.82) and reader expertise (OR, 2.06) were independently associated with adequate radiologist-LLM interaction. The confidence effect was weaker among expert readers (OR, 0.79), the study found. Also, higher rationale quality reduced readers’ rejection of correct suggestions (OR, 0.79) but increased their acceptance of incorrect suggestions (OR, 1.71). Higher reader expertise (OR, 0.54) and reader confidence (OR, 0.80) were protective, the study found, reducing rads’ acceptance of incorrect suggestions. 

“Beyond serving as a safeguard, reader expertise anchors human-LLM interaction,” the authors charged. “Given current limitations of vision-language models, expert-authored descriptions help guide model reasoning and reduce hallucinations. Expertise also enables clinicians to pose focused questions, evaluate outputs, and distinguish persuasive but unsound reasoning from clinically valid explanations. Rather than diminishing radiologists’ roles, LLMs amplify the importance of clinical expertise, as radiologists shape both inputs and interpretation.”

Read more, including potential study limitations, in the flagship journal of the Radiological Society of North America. 

Subscribe to Radiology Business News

Radiology Business Marty Stempniak

Marty Stempniak has covered healthcare since 2012, with his byline appearing in the American Hospital Association's member magazine, Modern Healthcare and McKnight's. Prior to that, he wrote about village government and local business for his hometown newspaper in Oak Park, Illinois. He won a Peter Lisagor and Gold EXCEL awards in 2017 for his coverage of the opioid epidemic. 

Subscribe to Radiology Business News

Subscribe to Radiology Business News