Enhanced Regularizers for Attributional Robustness
Anindya Sarkar, Anirban Sarkar, Vineeth N. Balasubramanian
Abstract
Deep neural networks are the default choice of learning models for computer vision tasks. Extensive work has been carried out in recent years on explaining deep models for vision tasks such as classification. However, recent work has shown that it is possible for these models to produce substantially different attribution maps even when two very similar images are given to the network, raising serious questions about trustworthiness. To address this issue, we propose a robust attribution training strategy to improve attributional robustness of deep neural networks. Our method carefully analyzes the requirements for attributional robustness and introduces two new regularizers that preserve a model's attribution map during attacks. Our method surpasses state-of-the-art attributional robustness methods by a margin of approximately 3% to 9% in terms of attribution robustness measures on several datasets including MNIST, FMNIST, Flower and GTSRB.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- Towards Automating Model Explanations with Certified Robustness GuaranteesMengdi Huai, Jinduo Liu, Chenglin Miao, Liuyi Yao et al.AAAI 2022 · 16 citations
- Exploiting the Relationship Between Kendall's Rank Correlation and Cosine Similarity for Attribution ProtectionFan Wang, Adams Wai-Kin KongNeurIPS 2022 · 12 citations
- Locally Invariant Explanations: Towards Stable and Unidirectional Explanations through Local Invariant LearningAmit Dhurandhar, Karthikeyan Natesan Ramamurthy, Kartik Ahuja, Vijay AryaNeurIPS 2023 · 7 citations
- Training for Stable Explanation for FreeChao Chen, Chenghua Guo, Rufeng Chen, Guixiang Ma et al.NeurIPS 2024 · 7 citations
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Relative Attributing Propagation: Interpreting the Comparative Contributions of Individual Units in Deep Neural NetworksWoo-Jeoung Nam, Shir Gur, Jaesik Choi, Lior Wolf et al.AAAI 2020 · 109 citations
- Jacobian Adversarially Regularized Networks for RobustnessAlvin Chan, Yi Tay, Yew-Soon Ong, Jie FuICLR 2020 · 81 citations
Related papers
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel et al.NeurIPS 2020 · 67 citations
- A Practical Upper Bound for the Worst-Case Attribution DeviationsFan Wang, Adams Wai-Kin KongCVPR 2023
- Improving Adversarial Robustness by Putting More Regularizations on Less Robust SamplesDongyoon Yang, Insung Kong, Yongdai KimICML 2023 · 15 citations
- AttEXplore: Attribution for Explanation with model parameters eXplorationZhiyu Zhu, Huaming Chen, Jiayu Zhang, Xinyi Wang et al.ICLR 2024 · 13 citations
- SAM: The Sensitivity of Attribution Methods to HyperparametersNaman Bansal, Chirag Agarwal, Anh NguyenCVPR 2020
