A Practical Upper Bound for the Worst-Case Attribution Deviations
Fan Wang, Adams Wai-Kin Kong
摘要
Model attribution is a critical component of deep neural networks (DNNs) for its interpretability to complex models. Recent studies bring up attention to the security of attribution methods as they are vulnerable to attribution attacks that generate similar images with dramatically different attributions. Existing works have been investigating empirically improving the robustness of DNNs against those attacks; however, none of them explicitly quantifies the actual deviations of attributions. In this work, for the first time, a constrained optimization problem is formulated to derive an upper bound that measures the largest dissimilarity of attributions after the samples are perturbed by any noises within a certain region while the classification results remain the same. Based on the formulation, different practical approaches are introduced to bound the attributions above using Euclidean distance and cosine similarity under both ℓ 2 and ℓ ∞ -norm perturbations constraints. The bounds developed by our theoretical study are validated on various datasets and two different types of attacks (PGD attack and IFIA attribution attack). Over 10 million attacks in the experiments indicate that the proposed upper bounds effectively quantify the robustness of models based on the worst-case attribution dissimilarities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 被引用 28 次
- Training for Stable Explanation for FreeChao Chen, Chenghua Guo, Rufeng Chen, Guixiang Ma 等NeurIPS 2024 · 被引用 7 次
- Pixel-level Certified Explanations via Randomized SmoothingAlaa Anani, Tobias Lorenz, Mario Fritz, Bernt SchieleICML 2025
它引用的顶会 Paper6
- How Does Mixup Help With Robustness and Generalization?Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani 等ICLR 2021 · 被引用 294 次
- Perceptual Adversarial Robustness: Defense Against Unseen Threat ModelsCassidy Laidlaw, Sahil Singla, Soheil FeiziICLR 2021 · 被引用 217 次
- Proper Network Interpretability Helps Adversarial Robustness in ClassificationAkhilan Boopathy, Sijia Liu, Gaoyuan Zhang, Cynthia Liu 等ICML 2020 · 被引用 74 次
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel 等NeurIPS 2020 · 被引用 67 次
- Enhanced Regularizers for Attributional RobustnessAnindya Sarkar, Anirban Sarkar, Vineeth N. BalasubramanianAAAI 2021 · 被引用 18 次
相关 Paper
- Towards More Robust Interpretation via Local Gradient AlignmentSunghwan Joo, Seokhyeon Jeong, Juyeon Heo, Adrian Weller 等AAAI 2023 · 被引用 8 次
- Exploiting the Relationship Between Kendall's Rank Correlation and Cosine Similarity for Attribution ProtectionFan Wang, Adams Wai-Kin KongNeurIPS 2022 · 被引用 12 次
- Manifold Integrated Gradients: Riemannian Geometry for Feature AttributionEslam Zaher, Maciej Trzaskowski, Quan Nguyen, Fred RoostaICML 2024 · 被引用 13 次
- AttEXplore: Attribution for Explanation with model parameters eXplorationZhiyu Zhu, Huaming Chen, Jiayu Zhang, Xinyi Wang 等ICLR 2024 · 被引用 13 次
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 被引用 480 次
