Robust Explanation Constraints for Neural Networks
Matthew Wicker, Juyeon Heo, Luca Costabello, Adrian Weller
摘要
Post-hoc explanation methods are used with the intent of providing insights about neural networks and are sometimes said to help engender trust in their outputs. However, popular explanations methods have been found to be fragile to minor perturbations of input features or model parameters. Relying on constraint relaxation techniques from non-convex optimization, we develop a method that upper-bounds the largest change an adversary can make to a gradient-based explanation via bounded manipulation of either the input features or model parameters. By propagating a compact input or parameter set as symbolic intervals through the forwards and backwards computations of the neural network we can formally certify the robustness of gradient-based explanations. Our bounds are differentiable, hence we can incorporate provable explanation robustness into neural network training. Empirically, our method surpasses the robustness provided by previous heuristic approaches. We find that our training method is the only method able to learn neural networks with certificates of explanation robustness across all six datasets tested.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SAFARI: Versatile and Efficient Evaluations for Robustness of InterpretabilityWei Huang, Xingyu Zhao, Gaojie Jin, Xiaowei HuangICCV 2023 · 被引用 41 次
- Learning with Explanation ConstraintsRattana Pukdee, Dylan Sam, J. Zico Kolter, Maria-Florina Balcan 等NeurIPS 2023 · 被引用 11 次
- Training for Stable Explanation for FreeChao Chen, Chenghua Guo, Rufeng Chen, Guixiang Ma 等NeurIPS 2024 · 被引用 7 次
- Certification for Differentially Private Prediction in Gradient-Based TrainingMatthew Wicker, Philip Sosnin, Igor Shilov, Adrianna Janik 等ICML 2025
它引用的顶会 Paper8
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller 等ICML 2020 · 被引用 104 次
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 被引用 93 次
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel 等NeurIPS 2020 · 被引用 67 次
相关 Paper
- Tight Certification of Adversarially Trained Neural Networks via Nonconvex Low-Rank Semidefinite RelaxationsHong-Ming Chiu, Richard Y. ZhangICML 2023 · 被引用 4 次
- Scalable Verified Training for Provably Robust Image ClassificationSven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel 等ICCV 2019 · 被引用 196 次
- Towards Better Understanding of Training Certifiably Robust Models against Adversarial ExamplesSungyoon Lee, Woojin Lee, Jinseong Park, Jaewook LeeNeurIPS 2021 · 被引用 27 次
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 被引用 28 次
- Interpreting Robustness Proofs of Deep Neural NetworksDebangshu Banerjee, Avaljot Singh, Gagandeep SinghICLR 2024 · 被引用 6 次
