U-CAM: Visual Explanation Using Uncertainty Based Class Activation Maps
Badri N. Patro, Mayank Lunayach, Shivansh Patel, Vinay P. Namboodiri
Abstract
Understanding and explaining deep learning models is an imperative task. Towards this, we propose a method that obtains gradient-based certainty estimates that also provide visual attention maps. Particularly, we solve for visual question answering task. We incorporate modern probabilistic deep learning methods that we further improve by using the gradients for these estimates. These have two-fold benefits: a) improvement in obtaining the certainty estimates that correlate better with misclassified samples and b) improved attention maps that provide state-of-the-art results in terms of correlation with human attention regions. The improved attention maps result in consistent improvement for various methods for visual question answering. Therefore, the proposed technique can be thought of as a recipe for obtaining improved certainty estimates and explanation for deep learning models. We provide detailed empirical analysis for the visual question answering task on all standard benchmarks and comparison with state of the art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 343f9d9f-844d-4907-8435-8f047c7726b5Cited by top-tier papers8
- Explanation vs Attention: A Two-Player Game to Obtain Attention for VQABadri N. Patro, Anupriy, Vinay P. NamboodiriAAAI 2020 · 27 citations
- Towards More Faithful Natural Language Explanation Using Multi-Level Contrastive Learning in VQAChengen Lai, Shengli Song, Shiqi Meng, Jingyang Li et al.AAAI 2024 · 12 citations
- Empowering CAM-Based Methods with Capability to Generate Fine-Grained and High-Faithfulness ExplanationsChangqing Qiu, Fusheng Jin, Yining ZhangAAAI 2024 · 11 citations
- Sentence Attention Blocks for Answer GroundingSeyedalireza Khoshsirat, Chandra KambhamettuICCV 2023 · 8 citations
- Flexible Visual Recognition by Evidential Modeling of Confusion and IgnoranceLei Fan, Bo Liu, Haoxiang Li, Ying Wu et al.ICCV 2023 · 7 citations
Related papers
- Re-Attention for Visual Question AnsweringWenya Guo, Ying Zhang, Xiaoping Wu, Jufeng Yang et al.AAAI 2020 · 90 citations
- VisQA: X-raying Vision and Language Reasoning in TransformersTheo Jaunet, Corentin Kervadec, Romain Vuillemot, Grigory Antipov et al.IEEE VIS 2021 · 30 citations
- SCOUT: Self-Aware Discriminant Counterfactual ExplanationsPei Wang, Nuno VasconcelosCVPR 2020
- Regularizing Attention Networks for Anomaly Detection in Visual Question AnsweringDoyup Lee, Yeongjae Cheon, Wook-Shin HanAAAI 2021 · 17 citations
- Gradient-based Uncertainty Attribution for Explainable Bayesian Deep LearningHanjing Wang, Dhiraj Joshi, Shiqiang Wang, Qiang JiCVPR 2023
