article · International Journal of Data Science and Analytics
Maintaining trust among employees, employers, and institutions is fundamental to business and research, yet the rise of always-online digital systems and expanding workforces poses new risks to operational integrity. Fluctuating work environments, evolving motivations, and gaps in training can leave organizations vulnerable to insider threats, including inadvertent data leaks and intentional exfiltration. In this study, an investigation of natural language processing (NLP) methods applied to HTTP activity logs to identify potential insider threats through behavioral patterns is conducted. Two experiments were conducted using TF-IDF and Word2Vec text representations, each combined with an XGBoost classifier whose hyperparameters were optimized using a newly proposed iteration stagnation-aware variable neighborhood search (ISAVNS) metaheuristic. The ISAVNS introduces a stagnation-detection mechanism that enables adaptive recovery during optimization, improving exploration and convergence stability. Evaluation on publicly available insider-threat datasets confirmed the high effectiveness of the proposed framework, with the TF-IDF-based model reaching an accuracy of 97.63% and the Word2Vec-based counterpart attaining 97.71%.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1007/s41060-025-00996-5
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.