MARATTO

article

Prosodic Parameterisation and Ensemble Machine Learning Techniques for Speech Emotion Detection

Abstract

This paper presents a robust approach to speech emotion recognition using an auto-correlated dynamic classical algorithm selection method. It employs FFT-based autocorrelation to extract fundamental frequency power spectral density for intensity, and inter-stress intervals for rhythm. Machine learning models (SVM. Decision Tree Naive Bayes, KNN, and Random Forest) were trained on RAVDESS, Emo-DB, TESS. IEMOCAP. and CREMA-D datasets, with ensemble learning achieving a state-of-the- art accuracy of 99.98%. The results highlight the effectiveness of the proposed approach in accurately detecting emotions from speech.

Research topics

  • Speech Recognition and Synthesis
  • Emotion and Mood Recognition
  • Speech and Audio Processing

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/nigercon62786.2024.10927084

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.