AI rubric could help make patient-friendly radiology reports safer
Artificial intelligence can translate complex radiology reports into language patients can more easily understand, but a new study suggests those AI-generated summaries may need their own quality control system before they are delivered to patients.
Researchers recently developed and evaluated a rubric designed to assess the quality and safety of AI-generated patient-friendly radiology reports. Their findings suggest the framework could provide a standardized safeguard for health systems incorporating AI into patient communications. The team shared their results Wednesday in the American Journal of Roentgenology.
The prospective study, which was conducted from February through December 2025, used a series of surveys and workshops involving lay participants and a multidisciplinary panel to develop the rubric. Researchers then used ChatGPT-4.1 and Claude-4 to generate patient-friendly versions of radiology report impressions from a public dataset. The reports were intentionally varied in quality across predefined criteria.
The resulting rubric evaluates five core characteristics—clarity, content, certainty, tone and verbosity, each of which is graded on a three-point scale. Under the rubric's decision rule, a report receiving the lowest grade for any attribute other than verbosity is considered unsafe or unacceptable for distribution to patients and should be withheld.
Researchers tested the rubric with radiologists, lay participants and an AI model. During development, three lay participants and three radiologists evaluating 60 reports demonstrated almost perfect agreement on overall grades. In broader testing involving 80 lay participants and 480 reports, agreement with prespecified reference standard grades was moderate. Subjective decisions about whether the reports were appropriate for distribution agreed with the rubric's rule-based decisions 73.5% of the time.
AI assessment of the same 480 reports also showed moderate agreement with the reference standard grades. When AI-assigned grades were used to determine whether reports should be distributed, the resulting decisions agreed with those based on reference standard grades 88.1% of the time.
The rubric is intended to address the patient communication gap by establishing explicit criteria for determining when an AI-generated report is acceptable for patient distribution and when it should be withheld for human review. Although additional training and validation are needed, the authors say rubric-based quality assurance has the potential to support more scalable and safer integration of AI-generated radiology communications into clinical workflows.
Read more here.
