Improving Interpretation Faithfulness for Vision Transformers
Lijie Hu, Yixin Liu, Ninghao Liu, Mengdi Huai, Lichao Sun, Di Wang
摘要
Vision Transformers (ViTs) have achieved state-of-the-art performance for various vision tasks. One reason behind the success lies in their ability to provide plausible innate explanations for the behavior of neural architectures. However, ViTs suffer from issues with explanation faithfulness, as their focal points are fragile to adversarial attacks and can be easily changed with even slight perturbations on the input image. In this paper, we propose a rigorous approach to mitigate these issues by introducing Faithful ViTs (FViTs). Briefly speaking, an FViT should have the following two properties: (1) The top- indices of its self-attention vector should remain mostly unchanged under input perturbation, indicating stable explanations; (2) The prediction distribution should be robust to perturbations. To achieve this, we propose a new method called Denoised Diffusion Smoothing (DDS), which adopts randomized smoothing and diffusion-based denoising. We theoretically prove that processing ViTs directly with DDS can turn them into FViTs. We also show that Gaussian noise is nearly optimal for both and -norm cases. Finally, we demonstrate the effectiveness of our approach through comprehensive experiments and evaluations. Results show that FViTs are more robust against adversarial attacks while maintaining the explainability of attention, indicating higher faithfulness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- SATO: Stable Text-to-Motion FrameworkWenshuo Chen, Hongru Xiao, Erhang Zhang, Lijie Hu 等ACM MM 2024 · 被引用 17 次
- Semi-Supervised Concept Bottleneck ModelsLijie Hu, Tianhao Huang, Huanyi Xie, Xilin Gong 等ICCV 2025 · 被引用 4 次
- Rethinking and Red-Teaming Protective Perturbation in Personalized Diffusion ModelsYixin Liu, Ruoxi Chen, Xun Chen, Lichao SunKDD 2026 · 被引用 3 次
- Benign Overfitting in Adversarial Training for Vision TransformersJiaming Zhang, Meng Ding, Shaopeng Fu, Jingfeng Zhang 等ICML 2026 · 被引用 1 次
- Vulcan: Crafting Compact Class-Specific Vision Transformers For Edge IntelligenceZiteng Wei, Qiang He, Feifei Chen, Ranjie Duan 等ICLR 2026
它引用的顶会 Paper18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
相关 Paper
- DAVE: Distribution-aware Attribution via ViT Gradient DecompositionAdam Wróbel, Siddhartha Gairola, Jacek Tabor, Bernt Schiele 等ICML 2026 · 被引用 2 次
- LeGrad: An Explainability Method for Vision Transformers via Feature Formation SensitivityWalid Bousselham, Angie W. Boggust, Sofian Chaybouti, Hendrik Strobelt 等ICCV 2025 · 被引用 47 次
- Manipulating the Mind's Eye: A-SAGE, the Attention-Based Attack on ViT ExplainabilityBoshi Zheng, Yan Li, Jiabin LiuAAAI 2026
- Towards Practical Certifiable Patch Defense with Vision TransformerZhaoyu Chen, Bo Li, Jianghe Xu, Shuang Wu 等CVPR 2022 · 被引用 60 次
- Denoising Diffusion Path: Attribution Noise Reduction with An Auxiliary Diffusion ModelYiming Lei, Zilong Li, Junping Zhang, Hongming ShanNeurIPS 2024 · 被引用 9 次
