Model confidence and reader expertise are independently associated with successful collaboration between radiologists and large language models in chest imaging, with higher expertise serving as a safeguard against accepting incorrect AI suggestions despite persuasive rationales.
- Model confidence (odds ratio 3.82) and reader expertise (odds ratio 2.06) were independently linked to effective reader-LLM interaction in chest imaging analysis.
- Higher quality LLM rationales reduced rejection of correct suggestions but increased acceptance of incorrect ones, creating a double-edged effect.
- Reader expertise and confidence helped reduce acceptance of incorrect suggestions, with expertise showing a stronger protective effect than confidence alone.
- Expert-authored text descriptions improved model accuracy from 27% to 76%, demonstrating how clinical expertise guides AI reasoning and reduces hallucinations.
Collaboration between readers and large language models (LLMs) may lead to improved model confidence and reader expertise in chest imaging, suggest findings published August 11 in Radiology.
The quality of rationales generated by LLMs leads to uptake of correct advice by readers but may also result in overreliance on incorrect suggestions by LLMs, wrote a team led by Jiyoung Song, MD, from Seoul National University Hospital in South Korea.
“Although high model confidence was associated with correct decisions, expertise may serve as a safeguard against persuasive but incorrect rationales,” the Song team wrote.
For radiologists, LLMs can provide diagnostic suggestions and rationales. However, the researchers noted a lack of data on what determines successful reader-LLM interaction.
Song and colleagues studied how LLM attributes and reader expertise influence how diagnostic advice is selectively integrated in human-LLM collaboration. For the retrospective study, 10 readers evaluated 100 chest imaging cases via x-ray, CT, MRI, or PET imaging.
The researchers implemented two LLM assistance setups for high and low LLM accuracies. The high-accuracy setup based on text input provided a highly capable reasoning model (ChatGPT-5, OpenAI) with structured clinical metadata and radiologist-authored image descriptions summarizing key findings as text inputs. GPT-5 achieved an accuracy of 76%. In the low-accuracy setup based on vision input, ChatGPT-4o (OpenAI) received the same metadata and the original chest images as visual inputs without text descriptions. This led to an accuracy of 27% for the model.
Example cases illustrate the double-edged role of large language model (LLM) rationale quality. (A) A 50-year-old female patient with a palpable anterior chest wall mass (correct diagnosis: solitary fibrous tumor of the pleura). The LLM incorrectly suggested synovial sarcoma but provided a persuasive rationale (quality score: 4 of 5), describing a well-defined nodule at axial contrast-enhanced CT spanning the pleura and intercostal space with heterogeneous strong enhancement. Four of five readers (80%) accepted the incorrect suggestion. FDG = fluorodeoxyglucose. (B) A 61-year-old male patient with diabetes presented with cough, sputum, and dyspnea (correct diagnosis: mucormycosis). The LLM correctly suggested mucormycosis with an excellent rationale (quality score: 5 of 5), integrating the reverse halo sign on serial axial CT images and rapidly progressive cavitary lesions. Three of five readers initially selected incorrect diagnoses; after LLM output, all five selected the correct diagnosis (100% acceptance). Blue and orange reader icons denote high- and low-expertise readers, respectively.RSNA
The team found after multivariable adjustment that model confidence (odds ratio [OR], 3.82; p = 0.003) and reader expertise (OR, 2.06; p < 0.001) were independently associated with adequate interaction between readers and LLMs. The researchers also noted a weaker effect of confidence among experts (OR, 0.79; p = 0.008).
Higher rationale quality reduced rejection of correct suggestions (OR, 0.79; p = 0.005) but increased acceptance of incorrect suggestions (OR, 1.71; p < 0.001).
Finally, the team reported that higher reader expertise (OR, 0.54; p < 0.001) and reader confidence (OR, 0.80; p = 0.007) were helped reduce the acceptance of incorrect suggestions.
The results point to effective LLM assistance going beyond model performance, the study authors wrote. They added that while reader expertise could serve as a safeguard in working with LLMs, it could also anchor interactions with LLMs.
“Given current limitations of vision-language models, expert-authored descriptions help guide model reasoning and reduce hallucinations,” the authors wrote. “Expertise also enables clinicians to pose focused questions, evaluate outputs, and distinguish persuasive but unsound reasoning from clinically valid explanations.”
They also noted that LLMs may amplify the importance of clinical expertise and that training “should evolve to develop skills in guiding model reasoning, interpreting outputs, and mitigating automation bias.”
Read the full study here.



















