Lune

NeurIPS2020顶会

Multimodal Graph Networks for Compositional Generalization in Visual Question Answering

Raeid Saqur, Karthik Narasimhan

出版方
2020年份
63被引次数
8顶会引用

摘要

Compositional generalization is a key challenge in grounding natural language to visual perception. While deep learning models have achieved great success in multimodal tasks like visual question answering, recent studies have shown that they fail to generalize to new inputs that are simply an unseen combination of those seen in the training distribution [6] . In this paper, we propose to tackle this challenge by employing neural factor graphs to induce a tighter coupling between concepts in different modalities (e.g. images and text). Graph representations are inherently compositional in nature and allow us to capture entities, attributes and relations in a scalable manner. Our model first creates a multimodal graph, processes it with a graph neural network to induce a factor correspondence matrix, and then outputs a symbolic program to predict answers to questions. Empirically, our model achieves close to perfect scores on a caption truth prediction problem and state-of-the-art results on the recently introduced CLOSURE dataset, improving on the mean overall accuracy across seven compositional templates by 4.77% over previous approaches. 2 * Work done at Princeton as a Fulbright Scholar. 2 Code is available at https://github.com/raeidsaqur/mgn 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada. individual nodes between modalities. A high-level overview of our approach is shown in Figure 1 , where one can observe fine-grained connections between nodes from graphs of both modalities. 0

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper8

问问它们各自怎么用它

它引用的顶会 Paper4

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖