EntropyStop: Unsupervised Deep Outlier Detection with Loss Entropy
Yihong Huang, Yuang Zhang, Liping Wang, Fan Zhang, Xuemin Lin
摘要
Unsupervised Outlier Detection (UOD) is an important data mining task. With the advance of deep learning, deep Outlier Detection (OD) has received broad interest. Most deep UOD models are trained exclusively on clean datasets to learn the distribution of the normal data, which requires huge manual efforts to clean the realworld data if possible. Instead of relying on clean datasets, some approaches directly train and detect on unlabeled contaminated datasets, leading to the need for methods that are robust to such challenging conditions. Ensemble methods emerged as a superior solution to enhance model robustness against contaminated training sets. However, the training time is greatly increased by the ensemble mechanism. In this study, we investigate the impact of outliers on training, aiming to halt training on unlabeled contaminated datasets before performance degradation. Initially, we noted that blending normal and anomalous data causes AUC fluctuations-a label-dependent measure of detection accuracy. To circumvent the need for labels, we propose a zero-label entropy metric named Loss Entropy for loss distribution, enabling us to infer optimal stopping points for training without labels. Meanwhile, a negative correlation between entropy metric and the label-based AUC score is demonstrated by theoretical proofs. Based on this, an automated early-stopping algorithm called EntropyStop is designed to halt training when loss entropy suggests the maximum model detection capability. We conduct extensive experiments on ADBench (including 47 real datasets), and the overall results indicate that AutoEncoder (AE) enhanced by our approach not only achieves better performance than ensemble AEs but also requires under 2% of training time. Lastly, loss entropy and EntropyStop are evaluated on other deep OD models, exhibiting their broad potential applicability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Unsupervised Anomaly Detection for Tabular Data Using Deep Noise EvaluationWei Dai, Kai Hwang, Jicong FanAAAI 2025 · 被引用 3 次
- AutoUAD: Hyper-parameter Optimization for Unsupervised Anomaly DetectionWei Dai, Jicong FanICLR 2025
它引用的顶会 Paper7
- Robust early-learning: Hindering the memorization of noisy labelsXiaobo Xia, Tongliang Liu, Bo Han, Chen Gong 等ICLR 2021 · 被引用 322 次
- Understanding and Improving Early Stopping for Learning with Noisy LabelsYingbin Bai, Erkun Yang, Bo Han, Yanhua Yang 等NeurIPS 2021 · 被引用 307 次
- Neural Transformation Learning for Deep Anomaly Detection Beyond ImagesChen Qiu, Timo Pfrommer, Marius Kloft, Stephan Mandt 等ICML 2021 · 被引用 171 次
- Anomaly Detection for Tabular Data with Internal Contrastive LearningTom Shenkar, Lior WolfICLR 2022 · 被引用 127 次
- InfoGAN-CR and ModelCentrality: Self-supervised Model Training and Selection for Disentangling GANsZinan Lin, Kiran Koshy Thekumparampil, Giulia Fanti, Sewoong OhICML 2020 · 被引用 106 次
相关 Paper
- Automatic Unsupervised Ensemble Outlier Model SelectionHong-Phuc Phan, Tuan-Anh Vu, Tung Kieu, Sơn Hà Xuân 等ICML 2026
- Hyperparameter Sensitivity in Deep Outlier Detection: Analysis and a Scalable Hyper-Ensemble SolutionXueying Ding, Lingxiao Zhao, Leman AkogluNeurIPS 2022 · 被引用 35 次
- Deep Semi-Supervised Anomaly DetectionLukas Ruff, Robert A. Vandermeulen, Nico Görnitz, Alexander Binder 等ICLR 2020 · 被引用 678 次
- An Encode-then-Decompose Approach to Unsupervised Time Series Anomaly Detection on Contaminated Training DataBuang Zhang, Tung Kieu, Xiangfei Qiu, Chenjuan Guo 等ICDE 2026 · 被引用 3 次
- Early Stopping Against Label Noise Without Validation DataSuqin Yuan, Lei Feng, Tongliang LiuICLR 2024 · 被引用 39 次
