MARATTO

article

Vision Language Transformation for Medical Image Captioning: Comparison of Four Pretrained CNN Networks

Abstract

Accurately interpreting medical images is crucial for effective clinical diagnoses, and the manual description of these images poses significant challenges. Medical image captioning bridges the gap between visual data and natural language, enabling the extraction of critical information from diverse medical datasets. This study introduces an innovative vision-language transformer approach that evaluates the performance of four powerful pre-trained convolutional neural networks (CNNs): InceptionV3, DenseNet121, ResNet50, and VGG16. By augmenting these networks with additional layers to enhance their trainability for feature extraction, we subsequently input the extracted features into a transformer model for caption generation. Our experiments utilized the Benchmark Radiography Captions (RGC) dataset, which encompasses various medical imaging modalities, including X-rays, MRIs, CT scans, and ultrasound images. To our knowledge, this is the first application of the RGC dataset for medical image captioning. We systematically compare the performance of the CNNs as feature extractors, assessing how well their output facilitates the transformer model's captioning capabilities. The source code for our approach is publicly available, promoting further research in the field of medical image captioning.

Research topics

  • Multimodal Machine Learning Applications
  • Image Retrieval and Classification Techniques
  • Advanced Image and Video Retrieval Techniques

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/iccta64612.2024.10974869

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.