article · Procedia Computer Science
Speech Emotion Recognition (SER) is a very interesting task that allows the machine to identify and recognize the different emotional states from human speech using new technologies. The SER can be represented by two main steps, namely feature extraction and emotion classification. Our contribution to the SER field will focus on these two phases. This paper seeks to investigate the influence of embedded features in Wav2vec2 and HuBERT models on SER, two variants per module are implemented, including Wav2vec2 base, Wav2vec2 large, HuBERT large and HuBERT X-large. In addition, we adopt a linear Support Vector Machine (SVM) as a downstream model to recognize emotions. The proposed approach relying on the combination of HuBERT X-large features with the SVM model led to the highest recognition rate of 82.6% on the RAVDESS database. Moreover, the results obtained are promising and in compliance with the current SER state-of-the-art. Hence, the embedded features of HuBERT X-large model have shown significant results for SER.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1016/j.procs.2024.02.074
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.