article
The exploration of emotion recognition from speech signals has garnered considerable attention across various domains. Instead of categorizing emotions into discrete classes such as anger and happiness, the dimensional continuous model is quite common. This study focuses on predicting arousal, valence, and dominance. Constructing a dimensional speech emotion recognition system involves selecting appropriate feature extraction and classification methods, a topic that remains under discussion. To tackle this challenge, this paper introduces a promising approach for extracting High Statistical Functions (HSFs) features from speech and evaluates two deep learning models for predicting arousal, valence, and dominance. The proposed models provide insights into the potential of optimized convolutional neural networks (CNN1D) and Long Short-Term Memory (LSTM) algorithms for emotion recognition tasks in speech signals. The optimized algorithms showed a significant result across both speaker-dependent and speaker-independent scenarios when evaluated on the IEMOCAP dataset.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/isivc61350.2024.10577881
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.