Explicit Knowledge Incorporation for Visual Reasoning
Yifeng Zhang, Ming Jiang, Qi Zhao
摘要
Existing explainable and explicit visual reasoning methods only perform reasoning based on visual evidence but do not take into account knowledge beyond what is in the visual scene. To addresses the knowledge gap between visual reasoning methods and the semantic complexity of realworld images, we present the first explicit visual reasoning method that incorporates external knowledge and models high-order relational attention for improved generalizability and explainability. Specifically, we propose a knowledge incorporation network that explicitly creates and includes new graph nodes for entities and predicates from external knowledge bases to enrich the semantics of the scene graph used in explicit reasoning. We then create a novel Graph-Relate module to perform high-order relational attention on the enriched scene graph. By explicitly introducing structured external knowledge and high-order relational attention, our method demonstrates significant generalizability and explainability over the state-of-the-art visual reasoning approaches on the GQA and VQAv2 datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada 等ICCV 2023 · 被引用 42 次
- MACK: Multimodal Aligned Conceptual Knowledge for Unpaired Image-text MatchingYan Huang, Yuming Wang, Yunan Zeng, Liang WangNeurIPS 2022 · 被引用 23 次
- Open-Vocabulary Object Detection With an Open CorpusJiong Wang, Huiming Zhang, Haiwen Hong, Xuan Jin 等ICCV 2023 · 被引用 22 次
- Query and Attention Augmentation for Knowledge-Based Explainable ReasoningYifeng Zhang, Ming Jiang, Qi ZhaoCVPR 2022 · 被引用 14 次
- Modality-Aware Integration with Large Language Models for Knowledge-Based Visual Question AnsweringJunnan Dong, Qinggang Zhang, Huachi Zhou, Daochen Zha 等ACL 2024 · 被引用 11 次
它引用的顶会 Paper1
相关 Paper
- From Strings to Things: Knowledge-Enabled VQA Model That Can Read and ReasonAjeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban ChakrabortyICCV 2019 · 被引用 54 次
- Let Me Show You Step by Step: An Interpretable Graph Routing Network for Knowledge-based Visual Question AnsweringDuokang Wang, Linmei Hu, Rui Hao, Yingxia Shao 等SIGIR 2024 · 被引用 2 次
- HAIR: Hierarchical Visual-Semantic Relational Reasoning for Video Question AnsweringFei Liu, Jing Liu, Weining Wang, Hanqing LuICCV 2021 · 被引用 58 次
- Scalable Multi-Hop Relational Reasoning for Knowledge-Aware Question AnsweringYanlin Feng, Xinyue Chen, Bill Yuchen Lin, Peifeng Wang 等EMNLP 2020 · 被引用 207 次
- Notes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question AnsweringWenlong Fang, Qiaofeng Wu, Jing Chen, Yun XueCVPR 2025
