MARATTO

article · Statistics, Optimization & Information Computing (University of Cambridge)

Decision-Level Fusion for Facial and Speech Emotion Recognition: A CNN-Based Web Application

2026Open accessCadi Ayyad University

Abstract

This paper presents a real-time web-based emotion recognition system based on unimodal deep learning models for facial and speech analysis, combined through decision-level score aggregation. Facial emotion recognition is performed using convolutional neural networks (CNNs), while speech emotion recognition relies on a CNN–BiLSTM architecture to capture both spatial and temporal speech patterns. These models are chosen for their effectiveness and low computational cost, making them suitable for web-based deployment. The facial model is trained on the FER2013 dataset, and the speech model is trained on the RAVDESS corpus using MFCC-based audio features. Rather than performing multimodal representation learning, this work demonstrates decision-level fusion by aggregating unimodal prediction scores to improve robustness when combining facial and speech information. Experimental results show competitive recognition performance and support the applicability of the proposed system for human-computer interaction in real-time and web-based affective applications.

Research topics

  • Emotion and Mood Recognition
  • Face and Expression Recognition
  • Face recognition and analysis

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.19139/soic-2310-5070-3341

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.