MARATTO

article · Scientific Reports

Optimized multi-modal action recognition using multi-agent systems and adaptive temporal attention

2026Open accessBeni Suef University

Abstract

The recognition of human actions through visual input poses considerable difficulties owing to the diverse ways in which individuals perform the same actions, the temporal variations that are intrinsic to these actions, and the differing viewpoints from which they are perceived. In order to address the limitations of single-modality methods, researchers have increasingly adopted multimodal visual data fusion strategies. This study presents an optimized deep learning model designed for action recognition by combining multiple data modalities, such as depth, skeleton, and inertial data through applying multi-agent systems. The study utilized a frame selection method based on the displacement of skeleton joints with uniform resampling for selection over the temporal-space; further, the study employed the HOG extractor on depth selected frames, 1D convolutional block for skeleton and inertial features extracting, and LSTM to process temporal patterns. Furthermore, an adaptive cross-attention method is developed to dynamically assess and implement the relative significance of the learned long-range temporal dependencies from each input modality. The effectiveness of the model is evaluated using the publicly available UTD MHAD dataset, demonstrating enhanced performance compared to existing leading action recognition techniques, achieving accuracy of 96.99%. A thorough assessment methodology, which includes confusion matrices, ROC curves, t-SNE visualizations, and average attention heatmaps, is utilized to confirm the model's robustness and to provide insights on its performance. The results underscore the model's capability in utilizing multi-modal data fusion and sophisticated temporal feature extraction, presenting a promising strategy for advancing action recognition tasks.

Research topics

  • Human Pose and Action Recognition
  • Hand Gesture Recognition Systems
  • Multimodal Machine Learning Applications

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1038/s41598-026-58924-x

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.