The downside of AI: Some rads accept recommendations even when they're incorrect
Radiologists who utilize artificial intelligence support may be vulnerable to automation bias, which can deteriorate their performance over time.
Automation bias refers to the tendency for physicians who use AI tools to over-rely on them. This can lead to radiologists accepting incorrect interpretations from AI and potentially spending less time evaluating exams than they would have without the added support. Though this is a known drawback of AI in medicine, the effects of automation bias have not been thoroughly researched.
A new paper in RSNA’s flagship journal Radiology delves deeper into how AI impacts radiologists’ reading habits and performance while interpreting screening mammograms. The authors describe a decrease in reader accuracy and visual search behaviors after rads were presented with an AI tool’s interpretation of an exam. This could be especially problematic when AI makes incorrect suggestions, experts caution.
“While the use of AI can improve workflows and efficiency and maintain good diagnostic performance, the risk of automation bias needs to be considered,” Paola Clauser, MD, PhD, with the department of biomedical imaging and image-guided therapy at the Medical University of Vienna, wrote in an editorial accompanying the study. “Current evidence suggests that less experienced readers are more prone to automation bias, while more experienced readers seem to be more capable of identifying wrong suggestions.”
Dr. Adnan G. Taib, with the University of Nottingham, and colleagues conducted a retrospective analysis of reader performance in the presence of AI to get a better idea of how these tools might impact radiologists performance. Ten breast radiologists were tasked with evaluating two-view mammography screening examinations twice, six weeks apart, both with and without the help of AI. Eye-tracking cameras were used during both interpretations to monitor how readers’ gaze changed for each test.
On the test set, the AI tool’s suggestions included a total of 26 true positives, 14 false negatives, 14 false positives and 6 true negatives.
Here’s how the readers fared:
Their median sensitivity for false negatives with AI was lower compared to their performance without AI, at 39% versus 71%.
Specificity was higher for cases with false positive AI suggestions (39% vs 21%).
True- and false-positive AI suggestions led to longer median read times, up from 25 seconds (with no prompts) to 34 seconds (with four or more prompts).
Compared to solo interpretations, readers tended to fixate less on cases the AI support deemed normal, although the exams actually had findings eventually diagnosed as malignant.
Shorter fixation durations also were observed when readers interpreted cases with false positive AI suggestions compared with unassisted reading.
The use of AI impacted both reader accuracy and behavior, especially in cases of false negative suggestions, which could be especially consequential for patients. The team suggested that their findings highlight the need for AI calibrations that could compensate for this accordingly.
The full study can be viewed here.
