Maintaining Reasoning Consistency in Compositional Visual Question Answering
Chenchen Jing, Yunde Jia, Yuwei Wu, Xinyu Liu, Qi Wu
Abstract
A compositional question refers to a question that contains multiple visual concepts (e.g., objects, attributes, and relationships) and requires compositional reasoning to answer. Existing VQA models can answer a compositional question well, but cannot work well in terms of reasoning consistency in answering the compositional question and its sub-questions. For example, a compositional question for an image is: "Are there any elephants to the right of the white bird?" and one of its sub-questions is " Is any bird visible in the scene?". The models may answer "yes" to the compositional question, but "no" to the sub-question. This paper presents a dialog-like reasoning method for maintaining reasoning consistency in answering a compositional question and its sub-questions. Our method integrates the reasoning processes for the sub-questions into the reasoning process for the compositional question like a dialog task, and uses a consistency constraint to penalize inconsistent answer predictions. In order to enable quantitative evaluation of reasoning consistency, we construct a GQA-Sub dataset based on the well-organized GQA dataset. Experimental results on the GQA dataset and the GQA-Sub dataset demonstrate the effectiveness of our method. * corresponding author (a) (b) Q: Is the train to the right or to the left of the black vehicle? (GT: left) Sub-Q1: Is there a vehicle that is black in this image? (GT: yes) Sub-Q2: What color does the vehicle have? (GT: black) left LCGN no green yes Sub-Q3: Do you see a train? (GT: yes) Q: Are there any elephants to the right of the white bird? (GT: yes)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 100cb33f-8ba6-4f07-a3fc-249727183d3bCited by top-tier papers8
- Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person RetrievalYiwei Ma, Xiaoshuai Sun, Jiayi Ji, Guannan Jiang et al.ACM MM 2023 · 34 citations
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter ItYulu Qin, Dheeraj Varghese, Adam Dahlgren Lindström, Lucia Donatelli et al.NeurIPS 2025 · 11 citations
- Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning ScenariosShantanu Jaiswal, Debaditya Roy, Basura Fernando, Cheston TanNeurIPS 2024 · 9 citations
- Detection-Based Intermediate Supervision for Visual Question AnsweringYuhang Liu, Daowan Peng, Wei Wei, Yuanyuan Fu et al.AAAI 2024 · 3 citations
- Core-to-Global Reasoning for Compositional Visual Question AnsweringHao Zhou, Tingjin Luo, Zhangqi JiangAAAI 2025 · 2 citations
Builds on7
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 183 citations
- Overcoming Language Priors in VQA via Decomposed Linguistic RepresentationsChenchen Jing, Yuwei Wu, Xiaoxun Zhang, Yunde Jia et al.AAAI 2020 · 115 citations
- Learning the Dynamics of Visual Relational Reasoning via Reinforced Path RoutingChenchen Jing, Yunde Jia, Yuwei Wu, Chuanhao Li et al.AAAI 2022 · 5 citations
- 12-in-1: Multi-Task Vision and Language Representation LearningJiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh et al.CVPR 2020
Related papers
- Consistency of Compositional Generalization Across Multiple LevelsChuanhao Li, Zhen Li, Chenchen Jing, Xiaomeng Fan et al.AAAI 2025 · 1 citation
- Measuring Compositional Consistency for Video Question AnsweringMona Gandhi, Mustafa Omer Gul, Eva Prakash, Madeleine Grunde-McLaughlin et al.CVPR 2022 · 11 citations
- Neuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning"Saeed Amizadeh, Hamid Palangi, Alex Polozov, Yichen Huang et al.ICML 2020 · 74 citations
- Interpretable Visual Reasoning via Induced Symbolic SpaceZhonghao Wang, Kai Wang, Mo Yu, Jinjun Xiong et al.ICCV 2021 · 22 citations
- Separating Skills and Concepts for Novel Visual Question AnsweringSpencer Whitehead, Hui Wu, Heng Ji, Rogério Feris et al.CVPR 2021
