article · Bulletin of Egyptian Society for Physiological Sciences
The exponential expansion of visual material in the digital era has driven developments in automatic image captioning to improve communication, accessibility, and visual content knowledge. This work presents a new encoder-decoder system using Xception network for feature extraction and a par-inject concatenate Recurrent Neural Network (RNN) for caption generation. Through a comprehensive comparison with Residual Network (ResNet-101)—the most commonly used encoder—we demonstrate the efficiency of the Xception encoder. Xception not only surpassed ResNet-101 by 22.8% in the Consensus Image Description Evaluation (CIDEr-D) metric but also outperformed it across all other evaluated metrics. Furthermore, Xception drastically decreases feature extraction time and memory use by 16% and 49%, respectively. The features extracted by Xception were 81.2 kilobytes while those extracted by Resnet101 were 165.9 kilobytes, indicating that the Xception performed the same task with almost half of image features detection required. This indicates that Xception is much more efficient in extracting image features. With interesting implications in assistive technologies and autonomous systems, this study presents a transforming method for image captioning that sets new standards for both accuracy and deployment efficiency.The efficiency of the framework and its outstanding performance make it ideal for real-time deployment on resource-constrained devices as smart glasses and mobile systems which can play back the textual captions as audio. This research focuses on deploying a system that can help visually impaired people perceive their surroundings. This assistive technology could completely change a visually impaired person’s day-to-day experience and help them live a more independent life.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.21608/besps.2025.398745.1222
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.