MARATTO

article

Integrating BERT and RESNET50V2 for Multimodal Cyberbullying Detection*

Abstract

Despite a range of advantages that social networks provide, cyberbullying is widespread. Several tools have been developed to automatically detect text-based cyberbullying, but with increased use of a combination of image and text in social media content, there is a need to detect cyberbullying from multimodal content such as images and text. In this paper, we present a custom architecture that uses two pretrained models for feature extraction, BERT and RESNET50V2. The proposed custom model architecture was trained on 149,823 instances of image-text pairs and produced a testing accuracy of 78%. Cross attention and use of dropouts were employed during model training, and focal loss was also used to make the model perform better on the rare class. A separate model that used state-of-the-art BERT for both feature extraction and classification of text only was trained on 149,823 instances of text and achieved an accuracy of 75%. The custom image and text model produced better results in comparison to state-of-the-art BERT for text, with its evaluation showing a good capability of detecting more cyberbullying in multimodal data.

Research topics

  • Hate Speech and Cyberbullying Detection
  • Bullying, Victimization, and Aggression
  • Advanced Malware Detection Techniques

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/icds62089.2024.10756337

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.