Bayesian Learning with Information Gain Provably Bounds Risk for a Robust Adversarial Defense
Bao Gia Doan, Ehsan Abbasnejad, Javen Qinfeng Shi, Damith C. Ranasinghe
Abstract
We present a new algorithm to learn a deep neural network model robust against adversarial attacks. Previous algorithms demonstrate an adversarially trained Bayesian Neural Network (BNN) provides improved robustness. We recognize the adversarial learning approach for approximating the multi-modal posterior distribution of a Bayesian model can lead to mode collapse; consequently, the model's achievements in robustness and performance are sub-optimal. Instead, we first propose preventing mode collapse to better approximate the multi-modal posterior distribution. Second, based on the intuition that a robust model should ignore perturbations and only consider the informative content of the input, we conceptualize and formulate an information gain objective to measure and force the information learned from both benign and adversarial training instances to be similar. Importantly. we prove and demonstrate that minimizing the information gain objective allows the adversarial risk to approach the conventional empirical risk. We believe our efforts provide a step toward a basis for a principled method of adversarially training BNNs. Our model demonstrate significantly improved robustness-up to 20%-compared with adversarial training (Madry et al., 2018) and Adv-BNN (Liu et al., 2019) under PGD attacks with 0.035 distortion on both CIFAR-10 and STL-10 datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 714fb5da-c464-4fd3-bf13-a288fc4fbb63Cited by top-tier papers3
- Bayesian Low-Rank Learning (Bella): A Practical Approach to Bayesian Neural NetworksBao Gia Doan, Afshar Shamsi, Xiao-Yu Guo, Arash Mohammadi et al.AAAI 2025 · 1 citation
- A Comprehensive Study of Deep Learning Model Fixing ApproachesHanmo You, Zan Wang, Zishuo Dong, Luanqi Mo et al.ICSE 2026
- Certified but Fooled! Breaking Certified Defenses with Ghost CertificatesViet Quoc Vo, Tashreque Mohammed Haq, Paul Montague, Tamas Abraham et al.AAAI 2026
Builds on7
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- HopSkipJumpAttack: A Query-Efficient Decision-Based AttackJianbo Chen, Michael I. Jordan, Martin J. WainwrightS&P 2020 · 797 citations
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 458 citations
- Sign-OPT: A Query-Efficient Hard-label Adversarial AttackMinhao Cheng, Simranjit Singh, Patrick H. Chen, Pin-Yu Chen et al.ICLR 2020 · 256 citations
- Robustness of Bayesian Neural Networks to Gradient-Based AttacksGinevra Carbone, Matthew Wicker, Luca Laurenti, Andrea Patané et al.NeurIPS 2020 · 85 citations
Related papers
- Improving Adversarial Robustness by Putting More Regularizations on Less Robust SamplesDongyoon Yang, Insung Kong, Yongdai KimICML 2023 · 15 citations
- Regularized Training and Tight Certification for Randomized Smoothed Classifier with Provable RobustnessHuijie Feng, Chunpeng Wu, Guoyang Chen, Weifeng Zhang et al.AAAI 2020 · 13 citations
- Improving Adversarial Robustness via Mutual Information EstimationDawei Zhou, Nannan Wang, Xinbo Gao, Bo Han et al.ICML 2022 · 23 citations
- Adversarial Training for Probabilistic RobustnessYi Zhang, Yuhang Chen, Zhen Chen, Wenjie Ruan et al.ICCV 2025 · 3 citations
- Splitting the Difference on Adversarial TrainingMatan Levi, Aryeh KontorovichUSENIX Security 2024 · 9 citations
