article
This article presents a state-of-the-art review of image captioning methodologies developed in the past five years. Image captioning, which aims to generate text describing the visual content of an image, has gained increasing interest in the field of artificial intelligence. The article analyzes how image captioning methods have evolved, transitioning from traditional machine learning-based approaches to deep learning-based methods that have become dominant due to their effectiveness in efficiently extracting image features. The article also explores two main approaches to image captioning: dense (region-based) captioning and whole-scene captioning. It discusses approaches based on visual space and multimodal space for text generation. Furthermore, the article examines reinforcement learning methods, semantic enhancements, the use of self-attention transformer models, and pretrained models such as the Generative Pre-trained Transformer (GPT) that have improved the performance of image captioning. Finally, the article highlights various applications of this technique, including human-machine interaction, biomedicine, automatic medical prescription, children’s education, industrial quality control, traffic data analysis, and assistive technologies for visually impaired individuals.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/wincom59760.2023.10322923
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.