Maintaining Reasoning Consistency in Compositional Visual Question Answering
Chenchen Jing, Yunde Jia, Yuwei Wu, Xinyu Liu, Qi Wu
摘要
A compositional question refers to a question that contains multiple visual concepts (e.g., objects, attributes, and relationships) and requires compositional reasoning to answer. Existing VQA models can answer a compositional question well, but cannot work well in terms of reasoning consistency in answering the compositional question and its sub-questions. For example, a compositional question for an image is: "Are there any elephants to the right of the white bird?" and one of its sub-questions is " Is any bird visible in the scene?". The models may answer "yes" to the compositional question, but "no" to the sub-question. This paper presents a dialog-like reasoning method for maintaining reasoning consistency in answering a compositional question and its sub-questions. Our method integrates the reasoning processes for the sub-questions into the reasoning process for the compositional question like a dialog task, and uses a consistency constraint to penalize inconsistent answer predictions. In order to enable quantitative evaluation of reasoning consistency, we construct a GQA-Sub dataset based on the well-organized GQA dataset. Experimental results on the GQA dataset and the GQA-Sub dataset demonstrate the effectiveness of our method. * corresponding author (a) (b) Q: Is the train to the right or to the left of the black vehicle? (GT: left) Sub-Q1: Is there a vehicle that is black in this image? (GT: yes) Sub-Q2: What color does the vehicle have? (GT: black) left LCGN no green yes Sub-Q3: Do you see a train? (GT: yes) Q: Are there any elephants to the right of the white bird? (GT: yes)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person RetrievalYiwei Ma, Xiaoshuai Sun, Jiayi Ji, Guannan Jiang 等ACM MM 2023 · 被引用 34 次
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter ItYulu Qin, Dheeraj Varghese, Adam Dahlgren Lindström, Lucia Donatelli 等NeurIPS 2025 · 被引用 11 次
- Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning ScenariosShantanu Jaiswal, Debaditya Roy, Basura Fernando, Cheston TanNeurIPS 2024 · 被引用 9 次
- Detection-Based Intermediate Supervision for Visual Question AnsweringYuhang Liu, Daowan Peng, Wei Wei, Yuanyuan Fu 等AAAI 2024 · 被引用 3 次
- Core-to-Global Reasoning for Compositional Visual Question AnsweringHao Zhou, Tingjin Luo, Zhangqi JiangAAAI 2025 · 被引用 2 次
它引用的顶会 Paper7
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 被引用 183 次
- Overcoming Language Priors in VQA via Decomposed Linguistic RepresentationsChenchen Jing, Yuwei Wu, Xiaoxun Zhang, Yunde Jia 等AAAI 2020 · 被引用 115 次
- Learning the Dynamics of Visual Relational Reasoning via Reinforced Path RoutingChenchen Jing, Yunde Jia, Yuwei Wu, Chuanhao Li 等AAAI 2022 · 被引用 5 次
- 12-in-1: Multi-Task Vision and Language Representation LearningJiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh 等CVPR 2020
相关 Paper
- Consistency of Compositional Generalization Across Multiple LevelsChuanhao Li, Zhen Li, Chenchen Jing, Xiaomeng Fan 等AAAI 2025 · 被引用 1 次
- Measuring Compositional Consistency for Video Question AnsweringMona Gandhi, Mustafa Omer Gul, Eva Prakash, Madeleine Grunde-McLaughlin 等CVPR 2022 · 被引用 11 次
- Neuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning"Saeed Amizadeh, Hamid Palangi, Alex Polozov, Yichen Huang 等ICML 2020 · 被引用 74 次
- Interpretable Visual Reasoning via Induced Symbolic SpaceZhonghao Wang, Kai Wang, Mo Yu, Jinjun Xiong 等ICCV 2021 · 被引用 22 次
- Separating Skills and Concepts for Novel Visual Question AnsweringSpencer Whitehead, Hui Wu, Heng Ji, Rogério Feris 等CVPR 2021
