Adversarial Robustness through Disentangled Representations
Shuo Yang, Tianyu Guo, Yunhe Wang, Chang Xu
摘要
Despite the remarkable empirical performance of deep learning models, their vulnerability to adversarial examples has been revealed in many studies. They are prone to make a susceptible prediction to the input with imperceptible adversarial perturbation. Although recent works have remarkably improved the model's robustness under the adversarial training strategy, an evident gap between the natural accuracy and adversarial robustness inevitably exists. In order to mitigate this problem, in this paper, we assume that the robust and non-robust representations are two basic ingredients entangled in the integral representation. For achieving adversarial robustness, the robust representations of natural and adversarial examples should be disentangled from the non-robust part and the alignment of the robust representations can bridge the gap between accuracy and robustness. Inspired by this motivation, we propose a novel defense method called Deep Robust Representation Disentanglement Network (DRRDN). Specifically, DRRDN employs a disentangler to extract and align the robust representations from both adversarial and natural examples. Theoretical analysis guarantees the mitigation of the trade-off between robustness and accuracy with good disentanglement and alignment performance. Experimental results on benchmark datasets finally demonstrate the empirical superiority of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Class-Disentanglement and Applications in Adversarial Detection and DefenseKaiwen Yang, Tianyi Zhou, Yonggang Zhang, Xinmei Tian 等NeurIPS 2021 · 被引用 49 次
- Combining Adversaries with Anti-adversaries in TrainingXiaoling Zhou, Nan Yang, Ou WuAAAI 2023 · 被引用 12 次
- CausalDiff: Causality-Inspired Disentanglement via Diffusion Model for Adversarial DefenseMingkun Zhang, Keping Bi, Wei Chen, Quanrun Chen 等NeurIPS 2024 · 被引用 9 次
- Mitigating Feature Gap for Adversarial Robustness by Feature DisentanglementNuoyan Zhou, Dawei Zhou, Decheng Liu, Nannan Wang 等AAAI 2025 · 被引用 3 次
- DifAttack: Query-Efficient Black-Box Adversarial Attack via Disentangled Feature SpaceJun Liu, Jiantao Zhou, Jiandian Zeng, Jinyu TianAAAI 2024 · 被引用 2 次
它引用的顶会 Paper1
相关 Paper
- Adversarial Invariant LearningNanyang Ye, Jingxuan Tang, Huayu Deng, Xiao-Yun Zhou 等CVPR 2021
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey 等ICLR 2020 · 被引用 829 次
- Splitting the Difference on Adversarial TrainingMatan Levi, Aryeh KontorovichUSENIX Security 2024 · 被引用 9 次
- Adversarial Finetuning with Latent Representation Constraint to Mitigate Accuracy-Robustness TradeoffSatoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Sekitoshi Kanai 等ICCV 2023 · 被引用 8 次
- Advancing Example Exploitation Can Alleviate Critical Challenges in Adversarial TrainingYao Ge, Yun Li, Keji Han, Junyi Zhu 等ICCV 2023 · 被引用 6 次
