article
Determining a person’s emotional state remains a non-trivial task relating to the ambiguous definition of the emotion itself and the different tools used to identify emotional aspects. In this study, we propose an Emotional Recognition System by adopting a multimodal approach that combines the based speech emotion features and the based facial expressions features. Accordingly, the proposed recognition system contains three parts. The first component will be reserved to extract the facial expressions features using a deep learning-based network called Long-Term Recurrent Convolutional Network (LRCN). The second component will be set aside to extract speech emotion features using the Alex-Net deep learning-based network. Finally, we dedicated the last component to combine the information learned from the two previous modalities through the late fusion. In order to evaluate the performance of the proposed approach, we tested it with the Ryerson Audio Visual Database of Emotional Speech and Song (RAVDESS) human emotions. Obtained results show that the accuracy rate was improved by using the fusion strategy. Indeed, the accuracy rate increased from 78.82 (for facial modality) and 76.39 (for speech modality) to 85.76.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1145/3607720.3607740
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.