Demystifying Causal Features on Adversarial Examples and Causal Inoculation for Robust Network by Adversarial Instrumental Variable Regression
Junho Kim, Byung-Kwan Lee, Yong Man Ro
摘要
The origin of adversarial examples is still inexplicable in research fields, and it arouses arguments from various viewpoints, albeit comprehensive investigations. In this paper, we propose a way of delving into the unexpected vulnerability in adversarially trained networks from a causal perspective, namely adversarial instrumental variable (IV) regression. By deploying it, we estimate the causal relation of adversarial prediction under an unbiased environment dissociated from unknown confounders. Our approach aims to demystify inherent causal features on adversarial examples by leveraging a zero-sum optimization game between a casual feature estimator (i.e., hypothesis model) and worstcase counterfactuals (i.e., test function) disturbing to find causal features. Through extensive analyses, we demonstrate that the estimated causal features are highly related to the correct prediction for adversarial robustness, and the counterfactuals exhibit extreme features significantly deviating from the correct prediction. In addition, we present how to effectively inoculate CAusal FEatures (CAFE) into defense networks for improving adversarial robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Meteor: Mamba-based Traversal of Rationale for Large Language and Vision ModelsByung-Kwan Lee, Chae Won Kim, Beomchan Park, Yong Man RoNeurIPS 2024 · 被引用 37 次
- Mitigating Adversarial Vulnerability through Causal Parameter Estimation by Adversarial Double Machine LearningByung-Kwan Lee, Junho Kim, Yong Man RoICCV 2023 · 被引用 12 次
- Unified Reinforcement and Imitation Learning for Vision-Language ModelsByung-Kwan Lee, Ryo Hachiuma, Yong Man Ro, Yu-Chiang Frank Wang 等NeurIPS 2025 · 被引用 12 次
- CausalDiff: Causality-Inspired Disentanglement via Diffusion Model for Adversarial DefenseMingkun Zhang, Keping Bi, Wei Chen, Quanrun Chen 等NeurIPS 2024 · 被引用 9 次
- TroL: Traversal of Layers for Large Language and Vision ModelsByung-Kwan Lee, Sangyun Chung, Chae Won Kim, Beomchan Park 等EMNLP 2024 · 被引用 5 次
它引用的顶会 Paper13
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
相关 Paper
- Adversarial Robustness Through the Lens of CausalityYonggang Zhang, Mingming Gong, Tongliang Liu, Gang Niu 等ICLR 2022 · 被引用 65 次
- Where and How to Attack? A Causality-Inspired Recipe for Generating Counterfactual Adversarial ExamplesRuichu Cai, Yuxuan Zhu, Jie Qiao, Zefeng Liang 等AAAI 2024 · 被引用 7 次
- Distribution-Conditioned Adversarial Variational Autoencoder for Valid Instrumental Variable GenerationXinshu Li, Lina YaoAAAI 2024 · 被引用 10 次
- Learning Deep Features in Instrumental Variable RegressionLiyuan Xu, Yutian Chen, Siddarth Srinivasan, Nando de Freitas 等ICLR 2021 · 被引用 85 次
- Data-faithful Feature Attribution: Mitigating Unobservable Confounders via Instrumental VariablesQiheng Sun, Haocheng Xia, Jinfei LiuNeurIPS 2024 · 被引用 3 次
