Separating Skills and Concepts for Novel Visual Question Answering
Spencer Whitehead, Hui Wu, Heng Ji, Rogério Feris, Kate Saenko
Abstract
Generalization to out-of-distribution data has been a problem for Visual Question Answering (VQA) models. To measure generalization to novel questions, we propose to separate them into "skills" and "concepts". "Skills" are visual tasks, such as counting or attribute recognition, and are applied to "concepts" mentioned in the question, such as objects and people. VQA methods should be able to compose skills and concepts in novel ways, regardless of whether the specific composition has been seen in training, yet we demonstrate that existing models have much to improve upon towards handling new compositions. We present a novel method for learning to compose skills and concepts that separates these two factors implicitly within a model by learning grounded concept representations and disentangling the encoding of skills from that of concepts. We enforce these properties with a novel contrastive learning procedure that does not rely on external annotations and can be learned from unlabeled image-question pairs. Experiments demonstrate the effectiveness of our approach for improving compositional and grounding performance. 1 * Work was partly done as an intern at the MIT-IBM Watson AI Lab.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb48e7a1-7205-4cb8-a1c8-e16a5608aa5eCited by top-tier papers14
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 732 citations
- DALL-EVAL: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation ModelsJaemin Cho, Abhay Zala, Mohit BansalICCV 2023 · 283 citations
- DUET: Cross-Modal Semantic Grounding for Contrastive Zero-Shot LearningZhuo Chen, Yufeng Huang, Jiaoyan Chen, Yuxia Geng et al.AAAI 2023 · 97 citations
- Modality-aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence DetectionJiashuo Yu, Jinyu Liu, Ying Cheng, Rui Feng et al.ACM MM 2022 · 64 citations
- Neural-Logic Human-Object Interaction DetectionLiulei Li, Jianan Wei, Wenguan Wang, Yi YangNeurIPS 2023 · 54 citations
Builds on6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin et al.ICCV 2019 · 288 citations
- Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption AlignmentSamyak Datta, Karan Sikka, Anirban Roy, Karuna Ahuja et al.ICCV 2019 · 113 citations
Related papers
- VQACL: A Novel Visual Question Answering Continual Learning SettingXi Zhang, Feifei Zhang, Changsheng XuCVPR 2023
- X-GGM: Graph Generative Modeling for Out-of-distribution Generalization in Visual Question AnsweringJingjing Jiang, Ziyi Liu, Yifan Liu, Zhixiong Nan et al.ACM MM 2021 · 17 citations
- Overcoming Language Priors in VQA via Decomposed Linguistic RepresentationsChenchen Jing, Yuwei Wu, Xiaoxun Zhang, Yunde Jia et al.AAAI 2020 · 115 citations
- Linguistically Routing Capsule Network for Out-of-distribution Visual Question AnsweringQingxing Cao, Wentao Wan, Keze Wang, Xiaodan Liang et al.ICCV 2021 · 16 citations
- Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation ModelsAmir Mohammad Karimi-Mamaghan, Samuele Papa, Karl Henrik Johansson, Stefan Bauer et al.ICLR 2025
