AttEXplore: Attribution for Explanation with model parameters eXploration
Zhiyu Zhu, Huaming Chen, Jiayu Zhang, Xinyi Wang, Zhibo Jin, Jason Xue, Flora D. Salim
Abstract
Due to the real-world noise and human-added perturbations, attaining the trustworthiness of deep neural networks (DNNs) is a challenging task. Therefore, it becomes essential to offer explanations for the decisions made by these nonlinear and complex parameterized models. Attribution methods are promising for this goal, yet its performance can be further improved. In this paper, for the first time, we present that the decision boundary exploration approaches of attribution are consistent with the process for transferable adversarial attacks. Specifically, the transferable adversarial attacks craft general adversarial samples from the source model, which is consistent with the generation of adversarial samples that can cross multiple decision boundaries in attribution. Utilizing this consistency, we introduce a novel attribution method via model parameter exploration. Furthermore, inspired by the capability of frequency exploration to investigate the model parameters, we provide enhanced explainability for DNNs by manipulating the input features based on frequency information to explore the decision boundaries of different models. Large-scale experiments demonstrate that our Attribution method for Explanation with model parameter eXploration (AttEXplore) outperforms other state-of-the-art interpretability methods. Moreover, by employing other transferable attack techniques, AttEXplore can explore potential variations in attribution outcomes. Our code is available at: https://github.com/LMBTough/ATTEXPLORE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e785702-5537-4cb4-a082-6448213ad76eCited by top-tier papers3
- Enhancing Model Interpretability with Local Attribution over Global ExplorationZhiyu Zhu, Zhibo Jin, Jiayu Zhang, Huaming ChenACM MM 2024 · 1 citation
- Narrowing Information Bottleneck Theory for Multimodal Image-Text Representations InterpretabilityZhiyu Zhu, Zhibo Jin, Jiayu Zhang, Nan Yang et al.ICLR 2025
- Faithfulness Under the Distribution: A New Look at Attribution EvaluationZhiyu Zhu, Zhibo Jin, Jiayu Zhang, Bartlomiej Sobieski et al.ICLR 2026
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang et al.ICLR 2020 · 765 citations
- Feature Importance-aware Transferable Adversarial AttacksZhibo Wang, Hengchang Guo, Zhifei Zhang, Wenxin Liu et al.ICCV 2021 · 306 citations
- Enhancing Adversarial Example Transferability With an Intermediate Level AttackQian Huang, Isay Katsman, Zeqi Gu, Horace He et al.ICCV 2019 · 293 citations
- Neural Pruning via Growing RegularizationHuan Wang, Can Qin, Yulun Zhang, Yun FuICLR 2021 · 188 citations
Related papers
- Smoothed Geometry for Robust AttributionZifan Wang, Haofan Wang, Shakul Ramkumar, Piotr Mardziel et al.NeurIPS 2020 · 67 citations
- Enhanced Regularizers for Attributional RobustnessAnindya Sarkar, Anirban Sarkar, Vineeth N. BalasubramanianAAAI 2021 · 18 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- Learning Transferable Adversarial PerturbationsKrishna Kanth Nakka, Mathieu SalzmannNeurIPS 2021 · 75 citations
- Towards Transferable Adversarial Attacks with Centralized PerturbationShangbo Wu, Yu-an Tan, Yajie Wang, Ruinan Ma et al.AAAI 2024 · 15 citations
