Multimodal Graph Networks for Compositional Generalization in Visual Question Answering
Raeid Saqur, Karthik Narasimhan
Abstract
Compositional generalization is a key challenge in grounding natural language to visual perception. While deep learning models have achieved great success in multimodal tasks like visual question answering, recent studies have shown that they fail to generalize to new inputs that are simply an unseen combination of those seen in the training distribution [6] . In this paper, we propose to tackle this challenge by employing neural factor graphs to induce a tighter coupling between concepts in different modalities (e.g. images and text). Graph representations are inherently compositional in nature and allow us to capture entities, attributes and relations in a scalable manner. Our model first creates a multimodal graph, processes it with a graph neural network to induce a factor correspondence matrix, and then outputs a symbolic program to predict answers to questions. Empirically, our model achieves close to perfect scores on a caption truth prediction problem and state-of-the-art results on the recently introduced CLOSURE dataset, improving on the mean overall accuracy across seven compositional templates by 4.77% over previous approaches. 2 * Work done at Princeton as a Fulbright Scholar. 2 Code is available at https://github.com/raeidsaqur/mgn 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada. individual nodes between modalities. A high-level overview of our approach is shown in Figure 1 , where one can observe fine-grained connections between nodes from graphs of both modalities. 0
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d70d1618-6091-4ed9-9d4b-68b1512eb3c2Cited by top-tier papers8
- Multi-level Feature Learning for Contrastive Multi-view ClusteringJie Xu, Huayi Tang, Yazhou Ren, Liang Peng et al.CVPR 2022 · 335 citations
- Debiased Visual Question Answering from Feature and Sample PerspectivesZhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu et al.NeurIPS 2021 · 102 citations
- Grammar-Based Grounded Lexicon LearningJiayuan Mao, Freda Shi, Jiajun Wu, Roger Levy et al.NeurIPS 2021 · 19 citations
- VisualHow: Multimodal Problem SolvingJinhui Yang, Xianyu Chen, Ming Jiang, Shi Chen et al.CVPR 2022 · 7 citations
- Improving compositional generalization for multi-step quantitative reasoning in question answeringArmineh Nourbakhsh, Cathy Jiao, Sameena Shah, Carolyn P. RoséEMNLP 2022 · 3 citations
Builds on4
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman et al.ICLR 2020 · 401 citations
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 183 citations
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt et al.NeurIPS 2020 · 169 citations
- Deep Graphical Feature Learning for the Feature Matching ProblemZhen Zhang, Wee Sun LeeICCV 2019 · 67 citations
Related papers
- Learning Visual Proxy for Compositional Zero-Shot LearningShiyu Zhang, Cheng Yan, Yang Liu, Chenchen Jing et al.ICCV 2025 · 1 citation
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu et al.CVPR 2022 · 63 citations
- Multi-modal Attentive Graph Pooling Model for Community Question Answer MatchingJun Hu, Quan Fang, Shengsheng Qian, Changsheng XuACM MM 2020 · 10 citations
- Learning to Represent Image and Text with Denotation GraphBowen Zhang, Hexiang Hu, Vihan Jain, Eugene Ie et al.EMNLP 2020 · 22 citations
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada et al.ICCV 2023 · 42 citations
