MARATTO

article

Improving Human Action Recognition in Videos with Two-Stream and Self-Attention Module

Abstract

Human action recognition has gained significant attention in recent years due to its potential applications in surveillance, human-computer interaction, healthcare, and security. Traditional approaches involve extracting low-level features from video sequences and using classifiers for labeling, while deep learning techniques have automated this process, providing more accurate and efficient action recognition. We propose a new approach for human action recognition in videos that employs a two-stream CNN to effectively capture appearance and motion information. Additionally, a self-attention network is used to selectively focus on the most informative regions of the video frames. Our method achieves an accuracy of 94.08% on the UCF101 dataset and demonstrates improved performance compared to existing techniques

Research topics

  • Human Pose and Action Recognition
  • Gait Recognition and Analysis
  • Hand Gesture Recognition Systems

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/cist56084.2023.10409877

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.