Training for Stable Explanation for Free
Chao Chen, Chenghua Guo, Rufeng Chen, Guixiang Ma, Ming Zeng, Xiangwen Liao, Xi Zhang, Sihong Xie
摘要
To foster trust in machine learning models, explanations must be faithful and stable for consistent insights. Existing relevant works rely on the ℓ p distance for stability assessment, which diverges from human perception. Besides, existing adversarial training (AT) associated with intensive computations may lead to an arms race. To address these challenges, we introduce a novel metric to assess the stability of top-k salient features. We introduce R2ET which trains for stable explanation by efficient and effective regularizer, and analyze R2ET by multi-objective optimization to prove numerical and statistical stability of explanations. Moreover, theoretical connections between R2ET and certified robustness justify R2ET’s stability in all attacks. Extensive experiments across various data modalities and model architectures show that R2ET achieves superior stability against stealthy attacks, and generalizes effectively across different explanation methods. The code can be found at https://github.com/ccha005/R2ET .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- On Gradient Descent Ascent for Nonconvex-Concave Minimax ProblemsTianyi Lin, Chi Jin, Michael I. JordanICML 2020 · 被引用 587 次
- On Explainability of Graph Neural Networks via Subgraph ExplorationsHao Yuan, Haiyang Yu, Jie Wang, Kang Li 等ICML 2021 · 被引用 498 次
- PGM-Explainer: Probabilistic Graphical Model Explanations for Graph Neural NetworksMinh N. Vu, My T. ThaiNeurIPS 2020 · 被引用 437 次
- A Closer Look at Accuracy vs. RobustnessYao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Ruslan Salakhutdinov 等NeurIPS 2020 · 被引用 336 次
相关 Paper
- Rethinking Robustness of Model AttributionsSandesh Kamath, Sankalp Mittal, Amit Deshpande, Vineeth N. BalasubramanianAAAI 2024 · 被引用 2 次
- Evaluations and Methods for Explanation through Robustness AnalysisCheng-Yu Hsieh, Chih-Kuan Yeh, Xuanqing Liu, Pradeep Kumar Ravikumar 等ICLR 2021 · 被引用 68 次
- Robust Explanation for Free or At the Cost of FaithfulnessZeren Tan, Yang TianICML 2023 · 被引用 12 次
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel 等NeurIPS 2020 · 被引用 67 次
- Can Adversarial Training Be Manipulated By Non-Robust Features?Lue Tao, Lei Feng, Hongxin Wei, Jinfeng Yi 等NeurIPS 2022 · 被引用 20 次
