MARATTO

preprint

Impact of Inaccurate Contamination Ratios on Robust Unsupervised Anomaly Detection: Experimental Investigation

Abstract

Training data sets intended for unsupervised anomaly detection, typically presumed to be anomaly-free, often contain anomalies (or contamination), a challenge that significantly undermines model performance. Most robust unsupervised anomaly detection models rely on contamination ratio information to tackle contamination. However, in reality, the contamination ratio may be inaccurate. We investigate the impact of inaccurate contamination ratio information in robust unsupervised anomaly detection. We verify whether they are resilient to misinformed contamination ratios. It appears that there is a lack of discussion in the literature regarding the impact of incorrect contamination ratio information on unsupervised anomaly detection, despite its critical importance. Our investigation on 8 benchmark data sets reveals that such models are not adversely affected by exposure to misinformation. In fact, they can exhibit improved performance when provided with such inaccurate contamination ratios.

Research topics

  • Anomaly Detection Techniques and Applications
  • Machine Learning and Data Classification
  • Water Systems and Optimization

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.36227/techrxiv.173396113.31607552/v1

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.