VLMs miss most lung nodules on chest x-rays

Vision-language AI models tested on lung nodule detection showed poor performance, with sensitivity ranging from less than 1% to 27%, indicating they are not ready for clinical use in chest x-ray interpretation despite earlier promising results.

  • Six vision-language models (VLMs) were tested on 247 chest x-rays containing lung nodules, with sensitivity ranging from less than 1% to 27%
  • LLaVA-Rad performed best with 27% sensitivity, while CheXpert Plus detected only one nodule out of 154 cases
  • Researchers concluded that zero-shot VLM performance does not support clinical use for lung nodule detection on radiographs
  • Study limitations included small sample size, single dataset, and lack of comparison with dedicated computer-aided detection systems
  • Future research may explore whether task-specific prompt engineering could improve nodule detection performance

Despite showing promise in early studies, recent findings do not support the use of vision-language AI models (VLMs) for detecting lung nodules on x-rays, according to a team in Japan. 

Researchers at Kobe University tested the “zero-shot” performance (no task-specific prompting or fine-tuning) of six VLMs on a dataset of 247 images, and although the models had moderate or even excellent specificity, sensitivity ranged from less than 1% to 27%, noted Mizuho Nishio, MD, PhD, and colleagues. 

“The findings therefore do not support use of the tested VLMs for lung nodule detection on radiographs and highlight the need for caution in their application for chest radiography interpretation,” the group wrote. The study was published August 5 in the American Journal of Roentgenology

Lung nodules are clinically important but commonly missed on x-rays, particularly when small or obscured by overlapping anatomy, the authors wrote. Although VLMs have shown ability to generate free-text interpretations of chest x-rays, their performance for flagging specific findings has remained poorly characterized, they added. 

To further test the models, the researchers used the Japanese Society of Radiological Technology (JSRT) chest radiograph database, which includes 247 posteroanterior chest radiographs: 154 containing a single nodule (mean size, 17.3 mm; confirmed by CT scans) and 93 without. The six VLMs tested were RadVLM, GPT-4o-mini, Qwen3-VL-8B-Instruct, MedGemma-4b-it, LLaVA-Rad, and CheXpert Plus. 

Five of the models received a generic prompt to generate a "concise report" with no nodule-specific instructions, while the CheXpert Plus received no text prompt at all. Two board-certified radiologists reviewed each model's output and classified it as positive or negative for nodule presence. 

LLaVA-Rad led on sensitivity at 27%. CheXpert Plus detected just one nodule in 154 cases. Full results appear in the table below.

Performance of VLMs for lung nodule detection on 247 chest radiographs 

VLM 

Sensitivity 

Specificity 

Accuracy 

RadVLM 

11% 

100% 

44.5% 

GPT-4o-mini 

6.5% 

93.5% 

39.3% 

Qwen3-VL-8B-Instruct 

7.8% 

88.2% 

38.1% 

MedGemma-4b-it 

5.2% 

100% 

40.9% 

LLaVA-Rad 

27.3% 

71% 

43.7% 

CheXpert Plus 

0.6% 

95.7% 

36.4% 

“This study’s findings indicate poor zero-shot performance of six VLMs for lung nodule detection on chest radiographs,” the authors wrote. 

Because of incomplete public documentation of the VLMs’ original training data, it could not be determined with certainty whether the VLMs had been exposed during development to the JSRT radiographs or to other resources developed using the JSRT radiographs, the group noted. 

Limitations of the study included the small sample size, use of a single dataset, lack of information regarding nodule type, incomplete public information regarding potential overlap between JSRT radiographs and VLM training data, lack of a human reader study, and lack of comparison with dedicated computer-aided detection systems, they added. 

“Future studies could evaluate whether task-specific prompt engineering improves nodule detection performance,” the authors concluded. 

The full study is available here, including an author video. 

Page 1 of 388
Next Page