MARATTO

article

Real-Time Speech Emotion Recognition

Abstract

This paper presents an analysis of three distinct speech emotion recognition methods, aiming to implement a real-time solution specifically for customer service applications in Morocco. The goal is to enhance the quality of customer relationship management in call centers by providing more convenient and usable emotion recognition tools. We used the CaFE speech dataset and we experimented with three methods. Namely, a Random Forest classifier with 100 decision trees, a three-layer Long Short-Term Memory (LSTM) model, and Multi-Layer Perceptron classifier combined with the state of the art pretrained speech model Wav2Vec for feature extraction. The results show that the best performing model is Random Forests with an accuracy of 85%. We integrate the model in a real-time scenario, showcasing its potential for practical applications in improving customer interactions in Moroccan call centers.

Research topics

  • Speech and Audio Processing

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/icds62089.2024.10756323

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.