article
The dawn of the Transformer era has advanced emotion recognition systems. Transformers capture contextual dependencies in data, making them highly effective for complex applications such as sentiment analysis and audio emotion recognition. In this study, we used self-supervised learning models such as wav2vec for speech and DistilBERT, a transformer-based model for text, to enhance the continuous emotion recognition system. These models are evaluated in the Interactive Emotional Dyadic Motion Capture (IEMOCAP) database, which contains speech data labeled with emotional dimensions such as valence, arousal, and dominance(VAD). Due to the complexity of the data and computational cost in fusing DistilBERT and wav2vec, we used an autoencoder for dimensionality reduction. To our knowledge, this is the first study to combine wav2vec and DistilBERT-like pre-trained features for Continuous Multimodal Emotion Recognition (CMER), addressing the challenge of limited labeled training data. Our experiments, evaluated using the Concordance Correlation Coefficient (CCC), show a significant performance boost, achieving a CCC of <tex>$0.808,0.719$</tex> and 0.635 respectively for VAD dimensions compared to <tex>$0.603,0.736$</tex> and 0.647 when using traditional feature extraction techniques (LSTM/CNN1D) on the IEMOCAP dataset. This demonstrates the effectiveness of using SSL models for emotion recognition tasks that typically suffer from small amounts of labeled data.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.23919/eusipco63237.2025.11226512
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.