Toward Multi-Granularity Decision-Making: Explicit Visual Reasoning with Hierarchical Knowledge
Yifeng Zhang, Shi Chen, Qi Zhao
摘要
Answering visual questions requires the ability to parse visual observations and correlate them with a variety of knowledge. Existing visual question answering (VQA) models either pay little attention to the role of knowledge or do not take into account the granularity of knowledge (e.g., attaching the color of "grassland" to "ground"). They have yet to develop the capability of modeling knowledge of multiple granularity, and are also vulnerable to spurious data biases. To fill the gap, this paper makes progresses from two distinct perspectives: (1) It presents a Hierarchical Concept Graph (HCG) that discriminates and associates multi-granularity concepts with a multi-layered hierarchical structure, aligning visual observations with knowledge across different levels to alleviate data biases. (2) To facilitate a comprehensive understanding of how knowledge contributes throughout the decision-making process, we further propose an interpretable Hierarchical Concept Neural Module Network (HCNMN). It explicitly propagates multi-granularity knowledge across the hierarchical structure and incorporates them with a sequence of reasoning steps, providing a transparent interface to elaborate on the integration of observations and knowledge. Through extensive experiments on multiple challenging datasets (i.e., GQA,VQA,FVQA,OK-VQA), we demonstrate the effectiveness of our method in answering questions in different scenarios. Our code is available at https://github.com/SuperJohnZhang/HCNMN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun 等AAAI 2021 · 被引用 414 次
- A Unified End-to-End Retriever-Reader Framework for Knowledge-based VQAYangyang Guo, Liqiang Nie, Yongkang Wong, Yibing Liu 等ACM MM 2022 · 被引用 41 次
- Explicit Knowledge Incorporation for Visual ReasoningYifeng Zhang, Ming Jiang, Qi ZhaoCVPR 2021
- KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQAKenneth Marino, Xinlei Chen, Devi Parikh, Abhinav Gupta 等CVPR 2021
相关 Paper
- Video as Conditional Graph Hierarchy for Multi-Granular Question AnsweringJunbin Xiao, Angela Yao, Zhiyuan Liu, Yicong Li 等AAAI 2022 · 被引用 145 次
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada 等ICCV 2023 · 被引用 42 次
- Progressive Graph Attention Network for Video Question AnsweringLiang Peng, Shuangji Yang, Yi Bin, Guoqing WangACM MM 2021 · 被引用 47 次
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai 等EMNLP 2020 · 被引用 157 次
- Let Me Show You Step by Step: An Interpretable Graph Routing Network for Knowledge-based Visual Question AnsweringDuokang Wang, Linmei Hu, Rui Hao, Yingxia Shao 等SIGIR 2024 · 被引用 2 次
