Demystifying Causal Features on Adversarial Examples and Causal Inoculation for Robust Network by Adversarial Instrumental Variable Regression
Junho Kim, Byung-Kwan Lee, Yong Man Ro
Abstract
The origin of adversarial examples is still inexplicable in research fields, and it arouses arguments from various viewpoints, albeit comprehensive investigations. In this paper, we propose a way of delving into the unexpected vulnerability in adversarially trained networks from a causal perspective, namely adversarial instrumental variable (IV) regression. By deploying it, we estimate the causal relation of adversarial prediction under an unbiased environment dissociated from unknown confounders. Our approach aims to demystify inherent causal features on adversarial examples by leveraging a zero-sum optimization game between a casual feature estimator (i.e., hypothesis model) and worstcase counterfactuals (i.e., test function) disturbing to find causal features. Through extensive analyses, we demonstrate that the estimated causal features are highly related to the correct prediction for adversarial robustness, and the counterfactuals exhibit extreme features significantly deviating from the correct prediction. In addition, we present how to effectively inoculate CAusal FEatures (CAFE) into defense networks for improving adversarial robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3074acf-8845-4ee3-9cdb-95a43a05db92Cited by top-tier papers10
- Meteor: Mamba-based Traversal of Rationale for Large Language and Vision ModelsByung-Kwan Lee, Chae Won Kim, Beomchan Park, Yong Man RoNeurIPS 2024 · 37 citations
- Mitigating Adversarial Vulnerability through Causal Parameter Estimation by Adversarial Double Machine LearningByung-Kwan Lee, Junho Kim, Yong Man RoICCV 2023 · 12 citations
- Unified Reinforcement and Imitation Learning for Vision-Language ModelsByung-Kwan Lee, Ryo Hachiuma, Yong Man Ro, Yu-Chiang Frank Wang et al.NeurIPS 2025 · 12 citations
- CausalDiff: Causality-Inspired Disentanglement via Diffusion Model for Adversarial DefenseMingkun Zhang, Keping Bi, Wei Chen, Quanrun Chen et al.NeurIPS 2024 · 9 citations
- TroL: Traversal of Layers for Large Language and Vision ModelsByung-Kwan Lee, Sangyun Chung, Chae Won Kim, Beomchan Park et al.EMNLP 2024 · 5 citations
Builds on13
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
Related papers
- Adversarial Robustness Through the Lens of CausalityYonggang Zhang, Mingming Gong, Tongliang Liu, Gang Niu et al.ICLR 2022 · 65 citations
- Where and How to Attack? A Causality-Inspired Recipe for Generating Counterfactual Adversarial ExamplesRuichu Cai, Yuxuan Zhu, Jie Qiao, Zefeng Liang et al.AAAI 2024 · 7 citations
- Distribution-Conditioned Adversarial Variational Autoencoder for Valid Instrumental Variable GenerationXinshu Li, Lina YaoAAAI 2024 · 10 citations
- Learning Deep Features in Instrumental Variable RegressionLiyuan Xu, Yutian Chen, Siddarth Srinivasan, Nando de Freitas et al.ICLR 2021 · 85 citations
- Data-faithful Feature Attribution: Mitigating Unobservable Confounders via Instrumental VariablesQiheng Sun, Haocheng Xia, Jinfei LiuNeurIPS 2024 · 3 citations
