Class-Disentanglement and Applications in Adversarial Detection and Defense
Kaiwen Yang, Tianyi Zhou, Yonggang Zhang, Xinmei Tian, Dacheng Tao
Abstract
What is the minimum necessary information required by a neural net D(•) from an image x to accurately predict its class? Extracting such information in the input space from x can allocate the areas D(•) mainly attending to and shed novel insights to the detection and defense of adversarial attacks. In this paper, we propose "class-disentanglement" that trains a variational autoencoder G(•) to extract this class-dependent information as x -G(x) via a trade-off between reconstructing x by G(x) and classifying x by D(x -G(x)), where the former competes with the latter in decomposing x so the latter retains only necessary information for classification in x -G(x). We apply it to both clean images and their adversarial images and discover that the perturbations generated by adversarial attacks mainly lie in the class-dependent part x -G(x). The decomposition results also provide novel interpretations to classification and attack models. Inspired by these observations, we propose to conduct adversarial detection and adversarial defense respectively on x -G(x) and G(x), which consistently outperform the results on the original x. In experiments, this simple approach substantially improves the detection and defense against different types of adversarial attacks. Code is available: https://github.com/kai-wen-yang/CD-VAE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 615fd63c-476e-48de-9d38-54d3dd393529Cited by top-tier papers12
- FedFed: Feature Distillation against Data Heterogeneity in Federated LearningZhiqin Yang, Yonggang Zhang, Yu Zheng, Xinmei Tian et al.NeurIPS 2023 · 166 citations
- Virtual Homogeneity Learning: Defending against Data Heterogeneity in Federated LearningZhenheng Tang, Yonggang Zhang, Shaohuai Shi, Xin He et al.ICML 2022 · 105 citations
- FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model FusionZhenheng Tang, Yonggang Zhang, Peijie Dong, Yiu-ming Cheung et al.NeurIPS 2024 · 28 citations
- Improving Adversarial Robustness via Mutual Information EstimationDawei Zhou, Nannan Wang, Xinbo Gao, Bo Han et al.ICML 2022 · 23 citations
- Adversarial Auto-Augment with Label Preservation: A Representation Learning Principle Guided ApproachKaiwen Yang, Yanchao Sun, Jiahao Su, Fengxiang He et al.NeurIPS 2022 · 14 citations
Builds on5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Self-Adaptive Training: beyond Empirical Risk MinimizationLang Huang, Chao Zhang, Hongyang ZhangNeurIPS 2020 · 256 citations
- Perceptual Adversarial Robustness: Defense Against Unseen Threat ModelsCassidy Laidlaw, Sahil Singla, Soheil FeiziICLR 2021 · 217 citations
- Adversarial Robustness through Disentangled RepresentationsShuo Yang, Tianyu Guo, Yunhe Wang, Chang XuAAAI 2021 · 38 citations
Related papers
- Improving VAEs' Robustness to Adversarial AttackMatthew Willetts, Alexander Camuto, Tom Rainforth, Stephen J. Roberts et al.ICLR 2021 · 30 citations
- Purify Unlearnable Examples via Rate-Constrained Variational AutoencodersYi Yu, Yufei Wang, Song Xia, Wenhan Yang et al.ICML 2024 · 22 citations
- Adversarial Defense by Restricting the Hidden Space of Deep Neural NetworksAamir Mustafa, Salman H. Khan, Munawar Hayat, Roland Goecke et al.ICCV 2019 · 160 citations
- Robustness and Generalization via Generative Adversarial TrainingOmid Poursaeed, Tianxing Jiang, Harry Yang, Serge J. Belongie et al.ICCV 2021 · 35 citations
- Delving into Deep Image Prior for Adversarial Defense: A Novel Reconstruction-based Defense FrameworkLi Ding, Yongwei Wang, Xin Ding, Kaiwen Yuan et al.ACM MM 2021 · 7 citations
