MARATTO

article

Multimodal Sarcasm Detection Method Using RNN and CNN

Abstract

In the current digital environment, sarcasm is preva-lent on social media platforms, characterized by a combination of verbal and non-verbal cues, such as prosodic variations, phonetic inflections, and textual markers including lexical selection, irony, and exaggeration. Previous research has primarily focused on sarcasm detection in either audio or text data independently. This paper presents a novel deep learning approach to detect sarcasm in conversational data by integrating both textual and auditory elements. Our method employs a bidirectional Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) network for text processing, capturing sequential dependencies and contextual information. For the audio component, we use a Convolutional Neural Network (CNN) to extract key features from speech. The fusion of these modalities is achieved by combining the extracted features into a composite vector, which improves the detection of sarcasm. Experimental evaluations on the MUStARD Extended dataset show that our hybrid model significantly outperforms unimodal models, achieving an F1-score of 74.67%.

Research topics

  • Face recognition and analysis

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/icds62089.2024.10756398

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.