Multimodal Graph Networks for Compositional Generalization in Visual Question Answering
Raeid Saqur, Karthik Narasimhan
摘要
Compositional generalization is a key challenge in grounding natural language to visual perception. While deep learning models have achieved great success in multimodal tasks like visual question answering, recent studies have shown that they fail to generalize to new inputs that are simply an unseen combination of those seen in the training distribution [6] . In this paper, we propose to tackle this challenge by employing neural factor graphs to induce a tighter coupling between concepts in different modalities (e.g. images and text). Graph representations are inherently compositional in nature and allow us to capture entities, attributes and relations in a scalable manner. Our model first creates a multimodal graph, processes it with a graph neural network to induce a factor correspondence matrix, and then outputs a symbolic program to predict answers to questions. Empirically, our model achieves close to perfect scores on a caption truth prediction problem and state-of-the-art results on the recently introduced CLOSURE dataset, improving on the mean overall accuracy across seven compositional templates by 4.77% over previous approaches. 2 * Work done at Princeton as a Fulbright Scholar. 2 Code is available at https://github.com/raeidsaqur/mgn 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada. individual nodes between modalities. A high-level overview of our approach is shown in Figure 1 , where one can observe fine-grained connections between nodes from graphs of both modalities. 0
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Multi-level Feature Learning for Contrastive Multi-view ClusteringJie Xu, Huayi Tang, Yazhou Ren, Liang Peng 等CVPR 2022 · 被引用 335 次
- Debiased Visual Question Answering from Feature and Sample PerspectivesZhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu 等NeurIPS 2021 · 被引用 102 次
- Grammar-Based Grounded Lexicon LearningJiayuan Mao, Freda Shi, Jiajun Wu, Roger Levy 等NeurIPS 2021 · 被引用 19 次
- VisualHow: Multimodal Problem SolvingJinhui Yang, Xianyu Chen, Ming Jiang, Shi Chen 等CVPR 2022 · 被引用 7 次
- Improving compositional generalization for multi-step quantitative reasoning in question answeringArmineh Nourbakhsh, Cathy Jiao, Sameena Shah, Carolyn P. RoséEMNLP 2022 · 被引用 3 次
它引用的顶会 Paper4
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman 等ICLR 2020 · 被引用 401 次
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 被引用 183 次
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt 等NeurIPS 2020 · 被引用 169 次
- Deep Graphical Feature Learning for the Feature Matching ProblemZhen Zhang, Wee Sun LeeICCV 2019 · 被引用 67 次
相关 Paper
- Learning Visual Proxy for Compositional Zero-Shot LearningShiyu Zhang, Cheng Yan, Yang Liu, Chenchen Jing 等ICCV 2025 · 被引用 1 次
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu 等CVPR 2022 · 被引用 63 次
- Multi-modal Attentive Graph Pooling Model for Community Question Answer MatchingJun Hu, Quan Fang, Shengsheng Qian, Changsheng XuACM MM 2020 · 被引用 10 次
- Learning to Represent Image and Text with Denotation GraphBowen Zhang, Hexiang Hu, Vihan Jain, Eugene Ie 等EMNLP 2020 · 被引用 22 次
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada 等ICCV 2023 · 被引用 42 次
