Leveraging Contaminated Datasets to Learn Clean-Data Distribution with Purified Generative Adversarial Networks
Bowen Tian, Qinliang Su, Jianxing Yu
Abstract
Generative adversarial networks (GANs) are known for their strong abilities on capturing the underlying distribution of training instances. Since the seminal work of GAN, many variants of GAN have been proposed. However, existing GANs are almost established on the assumption that the training dataset is clean. But in many real-world applications, this may not hold, that is, the training dataset may be contaminated by a proportion of undesired instances. When training on such datasets, existing GANs will learn a mixture distribution of desired and contaminated instances, rather than the desired distribution of desired data only (target distribution). To learn the target distribution from contaminated datasets, two purified generative adversarial networks (PuriGAN) are developed, in which the discriminators are augmented with the capability to distinguish between target and contaminated instances by leveraging an extra dataset solely composed of contamination instances. We prove that under some mild conditions, the proposed PuriGANs are guaranteed to converge to the distribution of desired instances. Experimental results on several datasets demonstrate that the proposed PuriGANs are able to generate much better images from the desired distribution than comparable baselines when trained on contaminated datasets. In addition, we also demonstrate the usefulness of PuriGAN on downstream applications by applying it to the tasks of semi-supervised anomaly detection on contaminated datasets and PU-learning. Experimental results show that PuriGAN is able to deliver the best performance over comparable baselines on both tasks 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a3e527e-db89-4e03-bd8b-80375f6bcbb6Cited by top-tier papers2
- Boosting Fine-Grained Visual Anomaly Detection with Coarse-Knowledge-Aware Adversarial LearningQingqing Fang, Qinliang Su, Wenxi Lv, Wenchao Xu et al.AAAI 2025 · 7 citations
- Contamination-Resilient Anomaly Detection via Adversarial Learning on Partially-Observed Normal and Anomalous DataWenxi Lv, Qinliang Su, Hai Wan, Hongteng Xu et al.ICML 2024 · 2 citations
Builds on6
- Deep Semi-Supervised Anomaly DetectionLukas Ruff, Robert A. Vandermeulen, Nico Görnitz, Alexander Binder et al.ICLR 2020 · 678 citations
- Gradient Normalization for Generative Adversarial NetworksYi-Lun Wu, Hong-Han Shuai, Zhi Rui Tam, Hong-Yu ChiuICCV 2021 · 78 citations
- Predictive Adversarial Learning from Positive and Unlabeled DataWenpeng Hu, Ran Le, Bing Liu, Feng Ji et al.AAAI 2021 · 56 citations
- Teaching a GAN What Not to LearnSiddarth Asokan, Chandra Sekhar SeelamantulaNeurIPS 2020 · 22 citations
- Negative Data AugmentationAbhishek Sinha, Kumar Ayush, Jiaming Song, Burak Uzkent et al.ICLR 2021 · 3 citations
Related papers
- GAN Ensemble for Anomaly DetectionXu Han, Xiaohui Chen, Li-Ping LiuAAAI 2021 · 79 citations
- Noise Robust Generative Adversarial NetworksTakuhiro Kaneko, Tatsuya HaradaCVPR 2020
- RGI: robust GAN-inversion for mask-free image inpainting and unsupervised pixel-wise anomaly detectionShancong Mou, Xiaoyi Gu, Meng Cao, Haoping Bai et al.ICLR 2023 · 8 citations
- AdaptiveMix: Improving GAN Training via Feature Space ShrinkageHaozhe Liu, Wentian Zhang, Bing Li, Haoqian Wu et al.CVPR 2023
- On Positive-Unlabeled Classification in GANTianyu Guo, Chang Xu, Jiajun Huang, Yunhe Wang et al.CVPR 2020
