Explanation vs Attention: A Two-Player Game to Obtain Attention for VQA
Badri N. Patro, Anupriy, Vinay P. Namboodiri
Abstract
In this paper, we aim to obtain improved attention for a visual question answering (VQA) task. It is challenging to provide supervision for attention. An observation we make is that visual explanations as obtained through class activation mappings (specifically Grad-CAM) that are meant to explain the performance of various networks could form a means of supervision. However, as the distributions of attention maps and that of Grad-CAMs differ, it would not be suitable to directly use these as a form of supervision. Rather, we propose the use of a discriminator that aims to distinguish samples of visual explanation and attention maps. The use of adversarial training of the attention regions as a two-player game between attention and explanation serves to bring the distributions of attention maps and visual explanations closer. Significantly, we observe that providing such a means of supervision also results in attention maps that are more closely related to human attention resulting in a substantial improvement over baseline stacked attention network (SAN) models. It also results in a good improvement in rank correlation metric on the VQA task. This method can also be combined with recent MCB based methods and results in consistent improvement. We also provide comparisons with other means for learning distributions such as based on Correlation Alignment (Coral), Maximum Mean Discrepancy (MMD) and Mean Square Error (MSE) losses and observe that the adversarial loss outperforms the other forms of learning the attention maps. Visualization of the results also confirms our hypothesis that attention maps improve using this form of supervision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80c832d8-5735-483f-9ee8-9a939cd63078Cited by top-tier papers2
- Attention-Aligned Transformer for Image CaptioningZhengcong FeiAAAI 2022 · 42 citations
- RES: A Robust Framework for Guiding Visual ExplanationYuyang Gao, Tong Steven Sun, Guangji Bai, Siyi Gu et al.KDD 2022 · 29 citations
Builds on1
Related papers
- QA-CLIMS: Question-Answer Cross Language Image Matching for Weakly Supervised Semantic SegmentationSonghe Deng, Wei Zhuo, Jinheng Xie, Linlin ShenACM MM 2023 · 12 citations
- Anti-Adversarially Manipulated Attributions for Weakly and Semi-Supervised Semantic SegmentationJungbeom Lee, Eunji Kim, Sungroh YoonCVPR 2021
- ODAM: Gradient-based Instance-Specific Visual Explanations for Object DetectionChenyang Zhao, Antoni B. ChanICLR 2023 · 4 citations
- Embedded Discriminative Attention Mechanism for Weakly Supervised Semantic SegmentationTong Wu, Junshi Huang, Guangyu Gao, Xiaoming Wei et al.CVPR 2021
- Towards Visually Explaining Variational AutoencodersWenQian Liu, Runze Li, Meng Zheng, Srikrishna Karanam et al.CVPR 2020
