Class-Disentanglement and Applications in Adversarial Detection and Defense
Kaiwen Yang, Tianyi Zhou, Yonggang Zhang, Xinmei Tian, Dacheng Tao
摘要
What is the minimum necessary information required by a neural net D(•) from an image x to accurately predict its class? Extracting such information in the input space from x can allocate the areas D(•) mainly attending to and shed novel insights to the detection and defense of adversarial attacks. In this paper, we propose "class-disentanglement" that trains a variational autoencoder G(•) to extract this class-dependent information as x -G(x) via a trade-off between reconstructing x by G(x) and classifying x by D(x -G(x)), where the former competes with the latter in decomposing x so the latter retains only necessary information for classification in x -G(x). We apply it to both clean images and their adversarial images and discover that the perturbations generated by adversarial attacks mainly lie in the class-dependent part x -G(x). The decomposition results also provide novel interpretations to classification and attack models. Inspired by these observations, we propose to conduct adversarial detection and adversarial defense respectively on x -G(x) and G(x), which consistently outperform the results on the original x. In experiments, this simple approach substantially improves the detection and defense against different types of adversarial attacks. Code is available: https://github.com/kai-wen-yang/CD-VAE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- FedFed: Feature Distillation against Data Heterogeneity in Federated LearningZhiqin Yang, Yonggang Zhang, Yu Zheng, Xinmei Tian 等NeurIPS 2023 · 被引用 166 次
- Virtual Homogeneity Learning: Defending against Data Heterogeneity in Federated LearningZhenheng Tang, Yonggang Zhang, Shaohuai Shi, Xin He 等ICML 2022 · 被引用 105 次
- FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model FusionZhenheng Tang, Yonggang Zhang, Peijie Dong, Yiu-ming Cheung 等NeurIPS 2024 · 被引用 28 次
- Improving Adversarial Robustness via Mutual Information EstimationDawei Zhou, Nannan Wang, Xinbo Gao, Bo Han 等ICML 2022 · 被引用 23 次
- Adversarial Auto-Augment with Label Preservation: A Representation Learning Principle Guided ApproachKaiwen Yang, Yanchao Sun, Jiahao Su, Fengxiang He 等NeurIPS 2022 · 被引用 14 次
它引用的顶会 Paper5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
- Self-Adaptive Training: beyond Empirical Risk MinimizationLang Huang, Chao Zhang, Hongyang ZhangNeurIPS 2020 · 被引用 256 次
- Perceptual Adversarial Robustness: Defense Against Unseen Threat ModelsCassidy Laidlaw, Sahil Singla, Soheil FeiziICLR 2021 · 被引用 217 次
- Adversarial Robustness through Disentangled RepresentationsShuo Yang, Tianyu Guo, Yunhe Wang, Chang XuAAAI 2021 · 被引用 38 次
相关 Paper
- Improving VAEs' Robustness to Adversarial AttackMatthew Willetts, Alexander Camuto, Tom Rainforth, Stephen J. Roberts 等ICLR 2021 · 被引用 30 次
- Purify Unlearnable Examples via Rate-Constrained Variational AutoencodersYi Yu, Yufei Wang, Song Xia, Wenhan Yang 等ICML 2024 · 被引用 22 次
- Adversarial Defense by Restricting the Hidden Space of Deep Neural NetworksAamir Mustafa, Salman H. Khan, Munawar Hayat, Roland Goecke 等ICCV 2019 · 被引用 160 次
- Robustness and Generalization via Generative Adversarial TrainingOmid Poursaeed, Tianxing Jiang, Harry Yang, Serge J. Belongie 等ICCV 2021 · 被引用 35 次
- Delving into Deep Image Prior for Adversarial Defense: A Novel Reconstruction-based Defense FrameworkLi Ding, Yongwei Wang, Xin Ding, Kaiwen Yuan 等ACM MM 2021 · 被引用 7 次
