Distilling Robust and Non-Robust Features in Adversarial Examples by Information Bottleneck
Junho Kim, Byung-Kwan Lee, Yong Man Ro
Abstract
Adversarial examples, generated by carefully crafted perturbation, have attracted considerable attention in research fields. Recent works have argued that the existence of the robust and non-robust features is a primary cause of the adversarial examples, and investigated their internal interactions in the feature space. In this paper, we propose a way of explicitly distilling feature representation into the robust and non-robust features, using Information Bottleneck. Specifically, we inject noise variation to each feature unit and evaluate the information flow in the feature representation to dichotomize feature units either robust or non-robust, based on the noise variation magnitude. Through comprehensive experiments, we demonstrate that the distilled features are highly correlated with adversarial prediction, and they have human-perceptible semantic information by themselves. Furthermore, we present an attack mechanism intensifying the gradient of non-robust features that is directly related to the model prediction, and validate its effectiveness of breaking model robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext adedad09-ddf9-414f-ac1a-247bb8f07337Cited by top-tier papers22
- Meteor: Mamba-based Traversal of Rationale for Large Language and Vision ModelsByung-Kwan Lee, Chae Won Kim, Beomchan Park, Yong Man RoNeurIPS 2024 · 37 citations
- What Can the Neural Tangent Kernel Tell Us About Adversarial Robustness?Nikolaos Tsilivis, Julia KempeNeurIPS 2022 · 28 citations
- Improving Adversarial Robustness via Information Bottleneck DistillationHuafeng Kuang, Hong Liu, Yongjian Wu, Shin'ichi Satoh et al.NeurIPS 2023 · 27 citations
- Can Adversarial Training Be Manipulated By Non-Robust Features?Lue Tao, Lei Feng, Hongxin Wei, Jinfeng Yi et al.NeurIPS 2022 · 20 citations
- Mitigating Adversarial Vulnerability through Causal Parameter Estimation by Adversarial Double Machine LearningByung-Kwan Lee, Junho Kim, Yong Man RoICCV 2023 · 12 citations
Builds on6
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Minimally distorted Adversarial Examples with a Fast Adaptive Boundary AttackFrancesco Croce, Matthias HeinICML 2020 · 597 citations
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 220 citations
Related papers
- Disentangled Information Bottleneck for Adversarial Text DefenseYidan Xu, Xinghao Yang, Wei Liu, Bao-di Liu et al.EMNLP 2025
- Attack to Explain Deep RepresentationMohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal MianCVPR 2020
- Mitigating Feature Gap for Adversarial Robustness by Feature DisentanglementNuoyan Zhou, Dawei Zhou, Decheng Liu, Nannan Wang et al.AAAI 2025 · 3 citations
- Fine-Grained Neural Network Explanation by Identifying Input Features with Predictive InformationYang Zhang, Ashkan Khakzar, Yawei Li, Azade Farshad et al.NeurIPS 2021 · 33 citations
- Feature Separation and Recalibration for Adversarial RobustnessWoo Jae Kim, Yoonki Cho, Junsik Jung, Sung-Eui YoonCVPR 2023
