VisFIS: Visual Feature Importance Supervision with Right-for-the-Right-Reason Objectives
Zhuofan Ying, Peter Hase, Mohit Bansal
Abstract
Many past works aim to improve visual reasoning in models by supervising feature importance (estimated by model explanation techniques) with human annotations such as highlights of important image regions. However, recent work has shown that performance gains from feature importance (FI) supervision for Visual Question Answering (VQA) tasks persist even with random supervision, suggesting that these methods do not meaningfully align model FI with human FI. In this paper, we show that model FI supervision can meaningfully improve VQA model accuracy as well as performance on several Right-for-the-Right-Reason (RRR) metrics by optimizing for four key model objectives: (1) accurate predictions given limited but sufficient information (Sufficiency); (2) max-entropy predictions given no important information (Uncertainty); (3) invariance of predictions to changes in unimportant features (Invariance); and (4) alignment between model FI explanations and human FI explanations (Plausibility). Our best performing method, Visual Feature Importance Supervision (VISFIS), outperforms strong baselines on benchmark VQA datasets in terms of both in-distribution and out-of-distribution accuracy. While past work suggests that the mechanism for improved accuracy is through improved explanation plausibility, we show that this relationship depends crucially on explanation faithfulness (whether explanations truly represent the model's internal reasoning). Predictions are more accurate when explanations are plausible and faithful, and not when they are plausible but not faithful. Lastly, we show that, surprisingly, RRR metrics are not predictive of out-of-distribution model accuracy when controlling for a model's in-distribution accuracy, which calls into question the value of these metrics for evaluating model reasoning. 2 * Equal contribution. 2 All supporting code for experiments in this paper is available at https://github.com/zfying/visfis .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c6d3dbc-4aa2-42db-88c5-2b4d21f955bbCited by top-tier papers3
- Adaptive Contextual Perception: How To Generalize To New Backgrounds and Ambiguous ObjectsZhuofan Ying, Peter Hase, Mohit BansalNeurIPS 2023 · 2 citations
- Uncovering the Full Potential of Visual Grounding Methods in VQADaniel Reich, Tanja SchultzACL 2024
- From Attribution to Action: Jointly ALIGNing Predictions and ExplanationsDongsheng Hong, Chao Chen, Yanhui Chen, Shanshan Lin et al.AAAI 2026
Builds on13
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini et al.NeurIPS 2020 · 731 citations
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin et al.ICCV 2019 · 288 citations
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha et al.NeurIPS 2020 · 163 citations
- The Effect of Natural Distribution Shift on Question Answering ModelsJohn Miller, Karl Krauth, Benjamin Recht, Ludwig SchmidtICML 2020 · 158 citations
- Improving Deep Learning Interpretability by Saliency Guided TrainingAya Abdelsalam Ismail, Héctor Corrada Bravo, Soheil FeiziNeurIPS 2021 · 121 citations
Related papers
- REX: Reasoning-aware and Grounded ExplanationShi Chen, Qi ZhaoCVPR 2022 · 24 citations
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question AnsweringZhihao Wen, Wenkang Wei, Yuan Fang, Xingtong Yu et al.CVPR 2026 · 1 citation
- RES: A Robust Framework for Guiding Visual ExplanationYuyang Gao, Tong Steven Sun, Guangji Bai, Siyi Gu et al.KDD 2022 · 29 citations
- Roses Are Red, Violets Are Blue... but Should VQA Expect Them To?Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian WolfCVPR 2021
- Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment RewardShizhan Gong, Minda Hu, Qiyuan Zhang, Chen Ma et al.CVPR 2026 · 1 citation
