MARATTO

article · International Journal of Intelligent Engineering Informatics

SwinRelTR: an efficient single-stage scene graph generation model for low-resolution images

20241 citationAin Shams University

Abstract

Targeting low-resolution imagery is crucial in democratising computer vision technologies, facilitating applicability in resource-limited environments where high-resolution data is often unreachable. Scene graphs have proven to be a powerful representation for capturing the hierarchical relationships between objects in an input image, providing a structured visual scene understanding. Nevertheless, all scene graph generation models focus on high-resolution images, neglecting the challenges posed by low-resolution images. This paper presents a novel approach called SwinRelTR for generating scene graphs designed specifically for low-resolution images. The proposed model addresses the limitations associated with low-resolution images by utilising the Swin transformer as a backbone instead of the convolution neural network in the original RelTR model. The Visual Genome dataset is utilised to compare the SwinRelTR results with the state-of-the-art approaches. It has been proven that this approach outperforms several state-of-the-art approaches as well as the original RelTR model on low-resolution images.

Research topics

  • Advanced Image and Video Retrieval Techniques
  • Advanced Neural Network Applications
  • Multimodal Machine Learning Applications

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1504/ijiei.2024.138854

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.