article
The increasing sophistication of cyber attacks necessitates effective intrusion detection systems. We propose a novel intrusion detection method integrating deep learning with big data management using Apache Spark. Leveraging the comprehensive CSE-CIC-IDS2018 dataset, we apply extensive data preprocessing, including handling missing and unreliable values, duplicates, and redundant columns. In addition, implementation of a Random Forest based feature importance approach is derived to prioritize the most impactful Features. Furthermore, stratified k-fold cross-validation is used for a model selection process on a class-imbalanced dataset. Our weighted Random Forest classifier achieves a remarkable weighted average F1-score of 0.999 and a test inference time of 0.673 seconds using only the top 34 features, outperforming previous studies without sampling techniques. The proposed architecture offers a scalable and accurate solution for intrusion detection in cloud architectures, demonstrating the effectiveness of combining deep learning and big data technologies for cybersecurity.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/niles63360.2024.10753188
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.