Multi-Modal Graph Neural Network for Joint Reasoning on Vision and Scene Text
Difei Gao, Ke Li, Ruiping Wang, Shiguang Shan, Xilin Chen
Abstract
Answering questions that require reading texts in an image is challenging for current models. One key difficulty of this task is that rare, polysemous, and ambiguous words frequently appear in images, e.g. names of places, products, and sports teams. To overcome this difficulty, only resorting to pre-trained word embedding models is far from enough. A desired model should utilize the rich information in multiple modalities of the image to help understand the meaning of scene texts, e.g. the prominent text on a bottle is most likely to be the brand. Following this idea, we propose a novel VQA approach, Multi-Modal Graph Neural Network (MM-GNN). It first represents an image as a graph consisting of three sub-graphs, depicting visual, semantic, and numeric modalities respectively. Then, we introduce three aggregators which guide the message passing from one graph to another to utilize the contexts in various modalities, so as to refine the features of nodes. The updated nodes have better features for the downstream question answering module. Experimental evaluations show that our MM-GNN represents the scene texts better and obviously facilitates the performances on two VQA tasks that require reading scene texts. * indicates equal contribution. A vision model can "see" Q4: Is the number in the image larger than 50? A: Yes Q3: How much is the super charge? A: 65 cents Q2: What color is the text on the top? A. Black Q1. What is the company who makes the product? A: STP A language model can "see" A calculator can "see" BERT Calculator A human can see (Original Image) Human CNN
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c7cdc18-55ed-41a9-b721-531f9f03d73dCited by top-tier papers20
- LaTr: Layout-Aware Transformer for Scene-Text VQAAli Furkan Biten, Ron Litman, Yusheng Xie, Srikar Appalaraju et al.CVPR 2022 · 82 citations
- Cascade Reasoning Network for Text-based Visual Question AnsweringFen Liu, Guanghui Xu, Qi Wu, Qing Du et al.ACM MM 2020 · 61 citations
- Simple is not Easy: A Simple Strong Baseline for TextVQA and TextCapsQi Zhu, Chenyu Gao, Peng Wang, Qi WuAAAI 2021 · 59 citations
- HAIR: Hierarchical Visual-Semantic Relational Reasoning for Video Question AnsweringFei Liu, Jing Liu, Weining Wang, Hanqing LuICCV 2021 · 58 citations
- Multimodal Continual Graph Learning with Neural Architecture SearchJie Cai, Xin Wang, Chaoyu Guan, Yateng Tang et al.WWW 2022 · 50 citations
Builds on2
Related papers
- From Strings to Things: Knowledge-Enabled VQA Model That Can Read and ReasonAjeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban ChakrabortyICCV 2019 · 54 citations
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada et al.ICCV 2023 · 42 citations
- Multi-modal Attentive Graph Pooling Model for Community Question Answer MatchingJun Hu, Quan Fang, Shengsheng Qian, Changsheng XuACM MM 2020 · 10 citations
- Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQAGangyan Zeng, Yuan Zhang, Yu Zhou, Xiaomeng YangACM MM 2021 · 38 citations
- Object Attribute Matters in Visual Question AnsweringPeize Li, Qingyi Si, Peng Fu, Zheng Lin et al.AAAI 2024 · 1 citation
