Stanford researchers developed a quality assessment rubric with five core attributes—clarity, content, certainty, tone, and verbosity—to safeguard AI-generated patient-friendly imaging reports before clinical release, achieving high agreement between lay audiences and radiologists while identifying areas for improvement.
- The rubric evaluates five core attributes: clarity, content, certainty, tone, and verbosity, each graded on a three-point scale
- Lay and radiologist evaluators achieved almost-perfect intergroup agreement (α=0.87) on overall report safety grades
- Rule-based distribution decisions aligned with subjective assessments in 91.2% of lay and 95.8% of radiologist evaluations
- The rubric addresses risks from AI translation errors that could compromise patient comprehension under the 21st Century Cures Act
- Further training and validation needed before clinical implementation, with ongoing rubric iterations recommended
A quality assessment rubric could serve as a safeguard before the release of patient-friendly reports generated by AI, suggest findings published September 9 in the American Journal of Roentgenology.
The rubric achieved high intergroup agreement between lay-friendly audiences and radiologists, wrote a team led by Bonnie Armstrong, PhD, from Stanford University in Palo Alto, California.
“Although requiring further training and validation, AI rubric application could enable scalable quality assurance and safer clinical integration of AI-generated communications,” the Armstrong team wrote.
Practices have begun to use large language models to generate patient-friendly imaging reports. This is in response to patients being able to access their reports (with technical language) before communicating with clinicians through the 21st Century Cures Act.
However, translation errors generated by these models could compromise patient comprehension and safety, Armstrong and colleagues noted. The researchers developed and evaluated a rubric that evaluates the quality and safety of patient-friendly radiology reports generated by AI.
They developed the rubric through survey-workshop cycles that involved lay participants and a multidisciplinary panel. From there, the researchers used ChatGPT-4.1 and Claude-4.0 to generate patient-friendly reports of varying quality based on radiology report impressions from a public dataset and prespecified quality targets across attributes.
For the prospective study, research team members, additional lay participants and radiologists, and ChatGPT-5 evaluated the generated reports using the rubric. In total, 19 participants developed the rubric, and six research team members and 111 more participants made up the evaluation part of the study.
The final rubric included five core attributes: clarity, content, certainty, tone, and verbosity. Each attribute was graded on a three-point scale. Patient-friendly reports assessed as grade 1 (unsafe or unacceptable) in any attribute other than verbosity are deemed unsafe to distribute to patients, according to the rubric’s rule.
The team reported the following findings:
Lay and radiologist research-team members (n=3 participants each; 60 reports) had almost-perfect intergroup agreement (α=0.87) for overall grade assignments.
Across six report pairs, overall interrater agreement across attributes was moderate (α = 0.51) among 19 lay participants and substantial (α = 0.65) among 12 radiologists.
The subjective distribution decisions among participants agreed with corresponding rule-derived distribution decisions based on participants’ assigned grades in 91.2% and 95.8% of lay participant and radiologist assessments, respectively.
In wider field testing, 80 lay participants (480 reports) had moderate agreement (κ = 0.43) with prespecified reference-standard grades. And subjective distribution decisions had a 73.5% agreement with rule-based distribution decisions.
Finally, across 480 reports, AI had moderate agreement (κ = 0.44) with prespecified reference-standard grades. Rule-based distribution decisions using AI-assigned grades had 88.1% agreement with rule-based distribution decisions using reference-standard grades.
The study authors highlighted the rubric being able to standardize otherwise subjective judgments of an AI-generated report’s patient-friendliness. Still, they noted the need for improvement for the rubric before it can be implemented into clinics.
“Agreement among raters and with the reference standard varied among attributes, and subjective distribution decisions occasionally deviated from rule-based decisions,” the authors wrote. “These observations provide feedback that could inform further rubric iterations.”
Read the full study here.



















