Toward Multi-Granularity Decision-Making: Explicit Visual Reasoning with Hierarchical Knowledge
Yifeng Zhang, Shi Chen, Qi Zhao
Abstract
Answering visual questions requires the ability to parse visual observations and correlate them with a variety of knowledge. Existing visual question answering (VQA) models either pay little attention to the role of knowledge or do not take into account the granularity of knowledge (e.g., attaching the color of "grassland" to "ground"). They have yet to develop the capability of modeling knowledge of multiple granularity, and are also vulnerable to spurious data biases. To fill the gap, this paper makes progresses from two distinct perspectives: (1) It presents a Hierarchical Concept Graph (HCG) that discriminates and associates multi-granularity concepts with a multi-layered hierarchical structure, aligning visual observations with knowledge across different levels to alleviate data biases. (2) To facilitate a comprehensive understanding of how knowledge contributes throughout the decision-making process, we further propose an interpretable Hierarchical Concept Neural Module Network (HCNMN). It explicitly propagates multi-granularity knowledge across the hierarchical structure and incorporates them with a sequence of reasoning steps, providing a transparent interface to elaborate on the integration of observations and knowledge. Through extensive experiments on multiple challenging datasets (i.e., GQA,VQA,FVQA,OK-VQA), we demonstrate the effectiveness of our method in answering questions in different scenarios. Our code is available at https://github.com/SuperJohnZhang/HCNMN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f9eaaeb-e249-4d07-a3ac-46a7a6157a69Cited by top-tier papers1
Ask how each one uses itBuilds on6
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun et al.AAAI 2021 · 414 citations
- A Unified End-to-End Retriever-Reader Framework for Knowledge-based VQAYangyang Guo, Liqiang Nie, Yongkang Wong, Yibing Liu et al.ACM MM 2022 · 41 citations
- Explicit Knowledge Incorporation for Visual ReasoningYifeng Zhang, Ming Jiang, Qi ZhaoCVPR 2021
- KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQAKenneth Marino, Xinlei Chen, Devi Parikh, Abhinav Gupta et al.CVPR 2021
Related papers
- Video as Conditional Graph Hierarchy for Multi-Granular Question AnsweringJunbin Xiao, Angela Yao, Zhiyuan Liu, Yicong Li et al.AAAI 2022 · 145 citations
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada et al.ICCV 2023 · 42 citations
- Progressive Graph Attention Network for Video Question AnsweringLiang Peng, Shuangji Yang, Yi Bin, Guoqing WangACM MM 2021 · 47 citations
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai et al.EMNLP 2020 · 157 citations
- Let Me Show You Step by Step: An Interpretable Graph Routing Network for Knowledge-based Visual Question AnsweringDuokang Wang, Linmei Hu, Rui Hao, Yingxia Shao et al.SIGIR 2024 · 2 citations
