LLM on par with head-and-neck subspecialists in analyzing MRI reports

Mri Brain Head 400

A study published in Radiology found that GPT-5-Thinking, a large language model, achieved diagnostic accuracy comparable to head-and-neck subspecialists when analyzing MRI reports of orbital and head-and-neck tumors, while also significantly improving diagnostic accuracy when used to assist generalist radiologists.

  • GPT-5-Thinking achieved 76.7% accuracy for orbital tumors and 67.5% for head-and-neck tumors, matching subspecialist performance (78.3% and 68.0% respectively)
  • When used as an aid, LLM assistance improved generalist accuracy by 8.9 percentage points for orbital tumors and 14 percentage points for head-and-neck tumors
  • The study included 1,000 patients with a mean age of 51 years, comparing LLM performance against routine clinical diagnoses and specialist interpretations
  • LLMs emulate specialist reasoning by analyzing growth patterns, enhancement kinetics, and clinical tempo to generate ranked differential diagnoses with supporting rationales
  • Researchers recommend integrating reasoning-oriented LLMs with structured reporting workflows while maintaining radiologist oversight for subspecialty-informed diagnosis support

Large language models (LLMs) may be on par with subspecialists in analyzing MRI reports of patients with orbital or head-and-neck tumors, according to research published September 15 in Radiology

GPT-5-Thinking achieved similar diagnostic accuracy compared to subspecialists while LLM assistance showed ties to higher generalist diagnostic accuracy, wrote a team led by Jie Li, MD, from The Second Hospital of Jilin University in China. 

“These findings suggest the potential of reasoning-oriented large language models as subspecialty-informed aids for generalist radiology practice,” Li and colleagues wrote. 

Researchers continue to explore how AI can assist imaging readers, including generalist readers. Li and co-authors noted that most deployed tools are image-centric algorithms that are limited to “narrow, task-specific” applications. 

LLMs meanwhile emulate semantic reasoning used by supspecialists, including analyzing cues such as growth patterns, enhancement kinetics, and clinical tempo. This these data, LLMs can generate a ranked differential with supporting rationales when providing a diagnosis from imaging reports. 

The Li team evaluated the performance of one such LLM, GPT-5-Thinking (OpenAI) and whether it could help bridge expertise gaps for generalist readers in interpreting MRI scans of orbital and head-and-neck tumors. 

Representative report-to-output workflow for GPT-5-Thinking interpretation of orbital MRI scan. This figure shows how deidentified clinical context and MRI report findings from one pretreatment patient-level report input were converted into a structured GPT-5-Thinking diagnostic output. (A) Images in a 55-year-old woman with a left orbital mass detected during physical examination. Axial fat-suppressed T2-weighted MRI scan shows a well-circumscribed hyperintense lesion in the left intraconal space posterior to the globe (arrow). Axial early and delayed postcontrast T1-weighted MRI scans show early peripheral nodular enhancement with progressive centripetal fill-in. (B) Fixed prompt used for GPT-5-Thinking interpretation. (C) Deidentified clinical history and MRI report findings used as model inputs. (D) Structured model output, including benign–malignant classification, ranked differential diagnosis, rationale, and conclusion.Representative report-to-output workflow for GPT-5-Thinking interpretation of orbital MRI scan. This figure shows how deidentified clinical context and MRI report findings from one pretreatment patient-level report input were converted into a structured GPT-5-Thinking diagnostic output. (A) Images in a 55-year-old woman with a left orbital mass detected during physical examination. Axial fat-suppressed T2-weighted MRI scan shows a well-circumscribed hyperintense lesion in the left intraconal space posterior to the globe (arrow). Axial early and delayed postcontrast T1-weighted MRI scans show early peripheral nodular enhancement with progressive centripetal fill-in. (B) Fixed prompt used for GPT-5-Thinking interpretation. (C) Deidentified clinical history and MRI report findings used as model inputs. (D) Structured model output, including benign–malignant classification, ranked differential diagnosis, rationale, and conclusion.RSNA

For the study, GPT-5-Thinking processed narrative MRI reports and generated classifications (benign or malignant) and top-three differential diagnoses. The researchers compared model accuracy with routine clinical report diagnoses in generalist and subspecialist practice settings. In one part of the study, seven generalists interpreted selected tumor reports twice, without and with GPT-5-Thinking aid. 

This study included 1000 patients (mean age ± SD, 51 years ± 17.0; 520 men).  

GPT-5-Thinking achieved a top-one accuracy similar to that of subspecialists in patients with orbital tumors (76.7% vs. 78.3%; p = 0.64) and head-and-neck tumors (67.5% vs. 68.0%; p = 0.92).  

GPT-5-Thinking also achieved higher top-one accuracy than routine generalists for patients with orbital tumors (74.0% vs. 68.7%; p = 0.03) and head-and-neck tumors (64.0% vs. 27.0%; p < 0.001).  

Assistance from the LLM also showed association with higher generalist top-one diagnostic accuracy. 

Comparison between assistance, no assistance with GPT-5-Thinking

Tumors

No LLM assistance

LLM assistance

P value

Orbital tumors

61.4%

70.3%

< 0.001

Head-and-neck tumors

47.0%

61.0%

< 0.001

LLMs such as GPT-5-Thinking could bridge gaps between integrating relevant signs and specific etiologic diagnoses, the study authors highlighted. This is due to LLMS being able to synthesize anatomic location, enhancement pattern, and disease-specific descriptors into a ranked differential diagnosis. 

“Future prospective studies should evaluate whether reasoning-oriented large language models can be integrated with structured reporting and image review workflows to support more consistent subspecialty-informed diagnosis while preserving radiologist oversight,” the authors wrote. 

Read the full study here.

Page 1 of 416
Next Page