GPT-4o outperforms radiologists at CT protocoling

Could large language models eventually automate the process of protocoling CT exams for radiologists? 

With proper prompt engineering, it is possible, according to new research published in RSNA’s flagship journal Radiology. However, context is the key to accurate and effective LLM protocoling, experts emphasize. 

“Accuracy is critical, as incorrect protocols can lead to nondiagnostic examinations, repeat imaging, and delayed diagnoses,” Rajesh Bhayana, MD, of University Medical Imaging Toronto, and colleagues noted. “However, protocoling is a manual and time-consuming task at most institutions that accounts for up to 6% of radiologists’ clinical time, competing with core interpretive responsibilities. Protocoling is also a recognized source of interruptions for radiologists, which can lead to increased diagnostic errors.” 

Context engineering has been shown to improve performance of LLMs, and newer models, like OpenAI’s GPT-4o, were designed to follow complex prompting instructions provided in context. Within the realm of CT scans of the abdomen and pelvis, context engineering could include information related to patients’ medical history, BMI, lab data, clinical complaints, scanner specs and more. 

Subscribe to Radiology Business News

Researchers recently examined how context engineering might affect an LLM’s ability to automatically assign protocols from CT scans of the abdomen and pelvis. To do this, they constructed detailed prompts and instructed GPT-4o to choose the most appropriate protocol from their institution’s list of 46 protocols that included detailed per-protocol selection criteria.  

The group tested the models on all abdomen and pelvis CT scans conducted at their facility between January and June of 2024. Data from the requisition and report, including the experience level of the provider who selected the protocol, was compared alongside the LLM’s outputs. 

With the help of context engineering, GPT-4o outperformed radiologists at selecting the most appropriate protocol, yielding an accuracy of 96.2% compared to the humans’ 88.3%. The team did not observe a difference between the LLM or humans in terms of inappropriate selections, and a subanalysis revealed that radiologists, fellows and residents all achieved similar results when choosing protocols matching the reference standard. 

“Our results suggest that for protocoling, state-of-the-art LLMs can be efficiently adapted with detailed prompt instructions to select optimal protocols more frequently than standard-of-care manual protocoling, without an increase in inappropriate protocols,” the group wrote. “Thus, LLMs could efficiently enable a more widespread use of automated protocol selection in supervised settings.” 

Read more here. 

Hannah Murphy
Hannah Murphy, Editor

In addition to her background in journalism, Hannah also has patient-facing experience in clinical settings, having spent more than 12 years working as a registered rad tech. She began covering the medical imaging industry for Innovate Healthcare in 2021.

Subscribe to Radiology Business News

Subscribe to Radiology Business News