article
In this paper we tackle the problem of violent event detection in realistic crowd scenes, which combine small scale individual actions with large scale group movement. We propose a dual-path motion encoding framework that jointly models fine-grained human micro-motions and global crowd dynamics. The local branch captures spatiotemporal human-centric features using 3D convolutions, while the global branch models scene-level temporal evolution through a bidirectional recurrent encoder. A stable training strategy combining warm-up scheduling and adaptive learning rate reduction ensures reliable convergence. Experimental results on the Real-Life Violence Situations dataset demonstrate that the proposed approach outperforms CNN, LSTM, and 3D CNN baselines in terms of accuracy and F1score. Visualization analysis confirms that the two branches learn complementary motion patterns. The method is suitable for realtime surveillance applications.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/iraset68627.2026.11538605
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.