Exploiting the Relationship Between Kendall's Rank Correlation and Cosine Similarity for Attribution Protection
Fan Wang, Adams Wai-Kin Kong
Abstract
Model attributions are important in deep neural networks as they aid practitioners in understanding the models, but recent studies reveal that attributions can be easily perturbed by adding imperceptible noise to the input. The non-differentiable Kendall's rank correlation is a key performance index for attribution protection. In this paper, we first show that the expected Kendall's rank correlation is positively correlated to cosine similarity and then indicate that the direction of attribution is the key to attribution robustness. Based on these findings, we explore the vector space of attribution to explain the shortcomings of attribution defense methods using norm and propose integrated gradient regularizer (IGR), which maximizes the cosine similarity between natural and perturbed attributions. Our analysis further exposes that IGR encourages neurons with the same activation states for natural samples and the corresponding perturbed samples, which is shown to induce robustness to gradient-based attribution methods. Our experiments on different models and datasets confirm our analysis on attribution protection and demonstrate a decent improvement in adversarial robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4bf5bee0-0ad2-4229-96da-f6b231f36cf2Cited by top-tier papers5
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- Manifold Integrated Gradients: Riemannian Geometry for Feature AttributionEslam Zaher, Maciej Trzaskowski, Quan Nguyen, Fred RoostaICML 2024 · 13 citations
- Training for Stable Explanation for FreeChao Chen, Chenghua Guo, Rufeng Chen, Guixiang Ma et al.NeurIPS 2024 · 7 citations
- Rethinking Robustness of Model AttributionsSandesh Kamath, Sankalp Mittal, Amit Deshpande, Vineeth N. BalasubramanianAAAI 2024 · 2 citations
- A Practical Upper Bound for the Worst-Case Attribution DeviationsFan Wang, Adams Wai-Kin KongCVPR 2023
Builds on7
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Proper Network Interpretability Helps Adversarial Robustness in ClassificationAkhilan Boopathy, Sijia Liu, Gaoyuan Zhang, Cynthia Liu et al.ICML 2020 · 74 citations
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel et al.NeurIPS 2020 · 67 citations
Related papers
- Towards More Robust Interpretation via Local Gradient AlignmentSunghwan Joo, Seokhyeon Jeong, Juyeon Heo, Adrian Weller et al.AAAI 2023 · 8 citations
- Enhanced Regularizers for Attributional RobustnessAnindya Sarkar, Anirban Sarkar, Vineeth N. BalasubramanianAAAI 2021 · 18 citations
- Attack-free Evaluating and Enhancing Adversarial Robustness on Categorical DataYujun Zhou, Yufei Han, Haomin Zhuang, Hongyan Bao et al.ICML 2024 · 2 citations
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu et al.ICML 2020 · 148 citations
- On the Robustness of Removal-Based Feature AttributionsChris Lin, Ian Covert, Su-In LeeNeurIPS 2023 · 25 citations
