Repository logo

CONTAMINATION-ROBUST AND DEPLOYMENT-ORIENTED UNSUPERVISED ANOMALY DETECTION UNDER UNKNOWN PREVALENCE AND STRUCTURED DATA DEFECTS

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Unsupervised anomaly detection is widely used in financial fraud detection, cybersecurity, IoT monitoring, and other settings where labeled anomalies are limited or unavailable. However, many anomaly detection methods are evaluated under assumptions that are difficult to satisfy in production, including clean and fully observed data, rare anomalies, known contamination levels, and stable score distributions. This thesis studies contamination-robust and deployment-oriented unsupervised anomaly detection under unknown anomaly prevalence and structured data defects. The thesis develops and evaluates a set of complementary methods and evaluation frameworks. First, UACE-IF-AD introduces unsupervised tail-mass calibration for Isolation Forest using density and tail diagnostics, including Gaussian mixture modeling, kernel density estimation, and extreme value modeling, to select auditable operating thresholds without labeled validation data or known contamination priors. Second, CAVET-AD extends the thresholding problem to detector-agnostic score calibration by mapping anomaly scores into a stabilized tail-calibrated space and studying threshold behavior under varying contamination regimes. Third, DRS-UAD introduces a defect-response evaluation framework that measures how anomaly detectors behave under controlled missingness and noise while preserving fixed train-test splits, label preservation, explicit failure logging, and fixed alert-rate evaluation. Fourth, CI-DCUAD investigates contamination-aware representation learning through reference selection, where a mask-aware encoder and contamination-informed subset selection are used to reduce the influence of contaminated samples during training. Finally, FM-UAD studies large-scale failure behavior under realistic data imperfections and increasing dataset size. Experiments are conducted across tabular financial fraud, network intrusion, and IoT security datasets, including CreditCard2013, NSL-KDD, CICIDS2018, IEEE-CIS, and CIC-IoT2023. The results show that threshold calibration can improve operating-point stability without changing the underlying score ranking, that clean-data performance alone is not sufficient to characterize deployment robustness, and that different detector families exhibit distinct degradation and failure patterns under missingness, noise, mixed corruption, and dataset scaling. The results also show that contamination-aware representation learning can improve robustness when score-space or embedding-space structure permits reliable reference selection, but its effectiveness is conditional and degrades when separability is weak. Overall, this thesis contributes a deployment-oriented view of unsupervised anomaly detection in which threshold calibration, defect-response behavior, execution viability, alert-budget control, and contamination-aware representation learning are studied jointly. The findings highlight the need to evaluate anomaly detection systems not only by clean benchmark accuracy, but also by their stability, auditability, and failure behavior under imperfect real-world data conditions.

Description

Citation

DOI

Collections

Endorsement

Review

Supplemented By

Referenced By