From Strings to Things: Knowledge-Enabled VQA Model That Can Read and Reason
Ajeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban Chakraborty
摘要
Text present in images are not merely strings, they provide useful cues about the image. Despite their utility in better image understanding, scene texts are not used in traditional visual question answering (VQA) models. In this work, we present a VQA model which can read scene texts and perform reasoning on a knowledge graph to arrive at an accurate answer. Our proposed model has three mutually interacting modules: (i) proposal module to get word and visual content proposals from the image, (ii) fusion module to fuse these proposals, question and knowledge base to mine relevant facts, and represent these facts as multi-relational graph, (iii) reasoning module to perform a novel gated graph neural network based reasoning on this graph. The performance of our knowledge-enabled VQA model is evaluated on our newly introduced dataset, viz. text-KVQA. To the best of our knowledge, this is the first dataset which identifies the need for bridging text recognition with knowledge graph based reasoning. Through extensive experiments, we show that our proposed method outperforms traditional VQA as well as question-answering over knowledge base-based methods on text-KVQA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Boosting Visual Question Answering with Context-aware Knowledge AggregationGuohao Li, Xin Wang, Wenwu ZhuACM MM 2020 · 被引用 82 次
- LaTr: Layout-Aware Transformer for Scene-Text VQAAli Furkan Biten, Ron Litman, Yusheng Xie, Srikar Appalaraju 等CVPR 2022 · 被引用 82 次
- TIVA-KG: A Multimodal Knowledge Graph with Text, Image, Video and AudioXin Wang, Benyuan Meng, Hong Chen, Yuan Meng 等ACM MM 2023 · 被引用 71 次
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada 等ICCV 2023 · 被引用 42 次
- Open-Vocabulary Object Detection With an Open CorpusJiong Wang, Huiming Zhang, Haiwen Hong, Xuan Jin 等ICCV 2023 · 被引用 22 次
它引用的顶会 Paper1
相关 Paper
- KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual DialogueXiaoze Jiang, Siyi Du, Zengchang Qin, Yajing Sun 等ACM MM 2020 · 被引用 37 次
- Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQAGangyan Zeng, Yuan Zhang, Yu Zhou, Xiaomeng YangACM MM 2021 · 被引用 38 次
- Iterative Answer Prediction With Pointer-Augmented Multimodal Transformers for TextVQARonghang Hu, Amanpreet Singh, Trevor Darrell, Marcus RohrbachCVPR 2020
- Cascade Reasoning Network for Text-based Visual Question AnsweringFen Liu, Guanghui Xu, Qi Wu, Qing Du 等ACM MM 2020 · 被引用 61 次
- Multi-Modal Graph Neural Network for Joint Reasoning on Vision and Scene TextDifei Gao, Ke Li, Ruiping Wang, Shiguang Shan 等CVPR 2020
