MARATTO

review · Applied Sciences

Multimodal Emotion Recognition Using Visual, Vocal and Physiological Signals: A Review

202442 citationsOpen accessTshwane University of Technology

In plain language

Automatically identifying human emotional states requires analysing various signals, including visual cues, vocal tones, and physiological data. Relying on a single modality often fails because specific noise factors interfere and distinct emotional states can appear identical. Integrating multiple modalities through deep learning improves accuracy and enables the detection of subtle micro-expressions across dynamic interactions. However, existing multimodal systems face significant operational hurdles. Key challenges include the scarcity of comprehensive training datasets, a lack of contextual awareness, and the degradation of performance caused by noisy or missing data in real-world settings. To address these vulnerabilities, future progress relies on developing enhanced input data representations, refining feature extraction techniques, and optimising how different sensory modalities are aggregated within unified computing frameworks.

Key takeaways

  • Single-modality emotion recognition systems are vulnerable to modality-specific noise and cannot always distinguish between different emotional states.
  • Combining visual, vocal, and physiological signals through multimodal fusion improves overall recognition accuracy and aids the detection of subtle micro-expressions.
  • Current deep learning methods for affective computing struggle with limited training data, poor context awareness, and real-world issues involving missing or noisy modalities.
  • Developing robust emotion recognition systems requires improved data representation, refined feature extraction, and optimised aggregation across modalities.

Why it matters

Human emotions are complex and dynamic, making automated understanding difficult when relying on a single channel like facial expressions or voice alone. By assessing how to combine multiple sensory inputs, this research highlights how machines can more reliably interpret human feelings and micro-expressions, even in noisy everyday conditions where certain sensory data might be incomplete or obscured.

Commercialisation angle

The abstract describes early-stage review research into affective computing frameworks. Such systems could ultimately enable developers and technology companies to create automated emotion recognition tools for human-computer interaction, health monitoring, or user experience evaluation. However, deployment remains distant from real-world market use due to ongoing challenges with limited training datasets, poor context awareness, and handling missing or noisy signals outside laboratory conditions.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

The dynamic expressions of emotion convey both the emotional and functional states of an individual’s interactions. Recognizing the emotional states helps us understand human feelings and thoughts. Systems and frameworks designed to recognize human emotional states automatically can use various affective signals as inputs, such as visual, vocal and physiological signals. However, emotion recognition via a single modality can be affected by various sources of noise that are specific to that modality and the fact that different emotion states may be indistinguishable. This review examines the current state of multimodal emotion recognition methods that integrate visual, vocal or physiological modalities for practical emotion computing. Recent empirical evidence on deep learning methods used for fine-grained recognition is reviewed, with discussions on the robustness issues of such methods. This review elaborates on the profound learning challenges and solutions required for a high-quality emotion recognition system, emphasizing the benefits of dynamic expression analysis, which aids in detecting subtle micro-expressions, and the importance of multimodal fusion for improving emotion recognition accuracy. The literature was comprehensively searched via databases with records covering the topic of affective computing, followed by rigorous screening and selection of relevant studies. The results show that the effectiveness of current multimodal emotion recognition methods is affected by the limited availability of training data, insufficient context awareness, and challenges posed by real-world cases of noisy or missing modalities. The findings suggest that improving emotion recognition requires better representation of input data, refined feature extraction, and optimized aggregation of modalities within a multimodal framework, along with incorporating state-of-the-art methods for recognizing dynamic expressions.

Research topics

  • Emotion and Mood Recognition
  • Color perception and design
  • Infant Health and Development

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.3390/app14178071

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.