A new study published in Radiology found that large language models can generate more comprehensive and clinically useful imaging indications than those provided by referring clinicians, potentially bridging a long-standing communication gap that contributes to diagnostic uncertainty and safety risks in radiology.
- LLM Performance: Claude 3.5 Sonnet achieved 37.14% top comprehensiveness ratings compared to 6.64% for referring clinicians, and 68.05% top factuality ratings versus 50% for clinicians.
- Study Scale: Researchers analyzed 740,867 clinical notes across 77,626 unique examinations from 28,313 patients at UCSF spanning January 2012 to August 2024.
- Clinical Validation: Twenty radiologists evaluated LLM-generated indications, rating them superior for protocoling, interpretation, and overall usefulness compared to original clinician indications.
- Persistent Problem: Inadequate imaging indications have persisted despite educational efforts, due to clinician awareness gaps and increasingly busy workflows that limit communication quality.
- Future Considerations: Experts emphasize transparency, continuous monitoring, and human oversight remain essential, with workflow efficiency and diagnostic improvements requiring further investigation.
Large language models (LLMs) can generate more comprehensive, clinically useful imaging indications than those provided by referring clinicians, according to a study published August 4 in Radiology.
The finding is from a test of open-source and proprietary LLMs on generating radiology-relevant clinical histories from patient electronic health records, and suggests LLMs can fill a communications gap with referring clinicians, noted lead authors Adrian Serapio, and Timothy L. Chen, MD, of the University of California, San Francisco (UCSF), and colleagues.
"Despite their importance, provided imaging indications are often inadequate partially due to the lack of awareness of their importance by ordering clinicians and increasingly busy clinician workflows," the group wrote. "Inadequate indications have persisted despite educational efforts directed at referring providers."
Substantial variability remains in clinicians' understanding of what constitutes radiologically relevant information, and the resultant communication gap is a well-recognized contributor to diagnostic uncertainty and avoidable safety risks, the group wrote. Faced with rising volumes and time pressure, radiologists may be forced to choose between proceeding with only the provided indication or restricting chart reviews to a few notes prone to inaccuracy, they added.
To determine whether LLMs could be useful in this setting, the researchers drew on deidentified UCSF electronic health records spanning January 2012 to August 2024, encompassing 28,313 patients (mean age, 59 years) and 740,867 clinical notes across 77,626 unique examinations. Each LLM received the referring clinician's original indication and the patient's 10 most recent clinical notes and generated an enhanced indication.
The group benchmarked multiple proprietary and open-source LLMs on automated performance metrics. The top proprietary model (Claude 3.5 Sonnet; Anthropic) and top open-source model (Qwen 2.5-7B Instruct; Alibaba) were then evaluated in a clinical reader study by 20 radiologists. Each reader compared 25 indications from the referring clinician and best-performing LLMs, scoring comprehensiveness, factuality, and conciseness, and ranking indications for usefulness in protocoling, usefulness in interpretation, and overall ranking.
Both LLMs outperformed referring clinicians on comprehensiveness and factuality. Claude 3.5 Sonnet received the highest percentage of top comprehensiveness ratings at 37.14%, followed by Qwen at 28.42% and referring clinicians at 6.64%. Top factuality ratings followed a similar pattern: 68.05% for Claude, 59.75% for Qwen, and 50% for referring clinicians.
In secondary results, comprehensiveness was the most commonly cited factor driving overall rankings, selected by 65.77% of readers. Claude 3.5 Sonnet received the highest proportion of top rankings for protocoling (40.87%), interpretation (44.61%), and overall (44.19%), outranking Qwen and the referring clinician in each category.
“This showcases the potential of LLMs beyond a simple summary of clinical notes, but also as a potential solution to address multiple persistent pain points in the radiologic workflow,” the researchers wrote.
In an accompanying editorial, Enis Yilmaz, MD, and David Cardoza-Ochoa, MD, both of the University of Texas Medical Branch, wrote that in line with the best-practice recommendations for deploying LLMs in radiology, transparency, continuous monitoring, and strong human oversight remain crucial.
“The results of this study suggest that LLMs can produce clinically relevant and comprehensive histories for radiologists,” they wrote.
Whether these benefits translate into measurable improvements in workflow efficiency, protocol selection, or diagnostic performance remains an important area for future investigation, Yilmaz, MD, and David Cardoza-Ochoa concluded.
The full study is available here.


















![A normal mammogram confirmed by three-year radiologic follow-up illustrates reader-marked regions of interest (ROIs) during (A) unaided (round 1) and (B) artificial intelligence (AI)–assisted (round 2) reading. Each colored dot represents an ROI for recall by a human reader. Readers could mark more than one ROI per case, represented by multiple dots of the same color. During AI-assisted reading, the AI system displayed three visible prompts: two with suspicion of malignancy scores of 35% (left mediolateral oblique [L MLO] and craniocaudal [L CC]) and one with a suspicion of malignancy score of 10% (right craniocaudal [R CC]), shown as polygonal overlays. Without AI, six of 10 readers (60%) marked a false-positive ROI. With AI assistance, this fell to two of 10 (20%). R MLO = right mediolateral oblique.](https://img.auntminnie.com/mindful/smg/workspaces/default/uploads/2026/07/2026-07-14-radiology-mammogram-ai-auto-bias.H0bYO8QlWs.jpg?auto=format%2Ccompress&fit=crop&h=112&q=70&w=112)
