article · Procedia Computer Science
In recent years, the prevalence of imbalanced datasets has posed significant challenges to traditional machine learning models. This imbalance is especially pronounced in fields such as spam detection, where malicious or unwanted messages are typically outnumbered by legitimate ones. Although various techniques have been developed to address this disparity, most conventional methods either undersample the majority class or oversample the minority class, potentially leading to information loss or overftting. In this study, we propose a novel approach using Generative Adversarial Networks (GANs) to generate synthetic samples, thus enhancing the representation of the minority class. By leveraging the powerful BERT embeddings to capture the intricate textual nuances, our model strives to produce synthetic spam messages that are not only realistic but also diverse. Initial results indicate that our GAN-augmented model offers a noticeable improvement in detecting spam messages compared to traditional techniques. This advancement not only holds potential for spam detection but also suggests broader applicability in addressing dataset imbalance across various domains.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1016/j.procs.2024.05.049
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.