Optimizing Relevance Maps of Vision Transformers Improves Robustness
Hila Chefer, Idan Schwartz, Lior Wolf
Abstract
It has been observed that visual classification models often rely mostly on the image background, neglecting the foreground, which hurts their robustness to distribution changes. To alleviate this shortcoming, we propose to monitor the model's relevancy signal and manipulate it such that the model is focused on the foreground object. This is done as a finetuning step, involving relatively few samples consisting of pairs of images and their associated foreground masks. Specifically, we encourage the model's relevancy map (i) to assign lower relevance to background regions, (ii) to consider as much information as possible from the foreground, and (iii) we encourage the decisions to have high confidence. When applied to Vision Transformer (ViT) models, a marked improvement in robustness to domain shifts is observed. Moreover, the foreground masks can be obtained automatically, from a self-supervised variant of the ViT model itself; therefore no additional supervision is required. Our code is available at https: //github.com/hila-chefer/RobustViT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4766319-dd1b-4ae2-81b4-90517c06ece1Cited by top-tier papers19
- Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion ModelsHila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf et al.SIGGRAPH 2023 · 438 citations
- Channel Vision Transformers: An Image Is Worth 1 x 16 x 16 WordsYujia Bao, Srinivasan Sivanandan, Theofanis KaraletsosICLR 2024 · 47 citations
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- Studying How to Efficiently and Effectively Guide Models with ExplanationsSukrut Rao, Moritz Böhle, Amin Parchami-Araghi, Bernt SchieleICCV 2023 · 22 citations
- Improving Accuracy-robustness Trade-off via Pixel Reweighted Adversarial TrainingJiacheng Zhang, Feng Liu, Dawei Zhou, Jingfeng Zhang et al.ICML 2024 · 9 citations
Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 451 citations
- Self-Supervised Transformers for Unsupervised Object Discovery using Normalized CutYangtao Wang, Xi Shen, Shell Xu Hu, Yuan Yuan et al.CVPR 2022 · 143 citations
Related papers
- Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve RobustnessYehonatan Elisha, Oren Barkan, Noam KoenigsteinCVPR 2026 · 2 citations
- Revisiting Continuity of Image Tokens for Cross-domain Few-shot LearningShuai Yi, Yixiong Zou, Yuhua Li, Ruixuan LiICML 2025
- Adapting Self-Supervised Vision Transformers by Probing Attention-Conditioned Masking ConsistencyViraj Prabhu, Sriram Yenamandra, Aaditya Singh, Judy HoffmanNeurIPS 2022 · 17 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- Patch-level Representation Learning for Self-supervised Vision TransformersSukmin Yun, Hankook Lee, Jaehyung Kim, Jinwoo ShinCVPR 2022 · 52 citations
