MARATTO

article

Hybrid Architecture and pre-processing for Speech Emotion Recognition

Abstract

Accurate emotion recognition from speech is vital for improving human-computer interaction, particularly in advanced Mechatronic systems where responsive and adaptive behaviour is crucial. Traditional deep learning models like CNNs and LSTMs often need help with transient emotional nuances, and the presence of temporal relationships in audio signals limits their effectiveness in practical scenarios. This study introduces the Dendritic Convolutional LSTM (DCLSTM) architecture, which integrates dendritic computational principles to enhance the learning of complex spatial and temporal features in emotional speech. An advanced audio preprocessing pipeline is also implemented, systematically refining speech signals through noise reduction, spectral processing, and filtering to optimise model performance. By enabling more accurate and nuanced emotion recognition, this research represents a significant advancement in integrating human-like emotional understanding into Mechatronic systems, leading to greater effortless and flexible machine interactions.

Research topics

  • Speech Recognition and Synthesis
  • Speech and Audio Processing
  • Emotion and Mood Recognition

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/icamechs63130.2024.10818817

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.