Robust Visual Reasoning via Language Guided Neural Module Networks
Arjun R. Akula, Varun Jampani, Soravit Changpinyo, Song-Chun Zhu
摘要
Neural module networks (NMN) are a popular approach for solving multi-modal tasks such as visual question answering (VQA) and visual referring expression recognition (REF). A key limitation in prior implementations of NMN is that the neural modules do not effectively capture the association between the visual input and the relevant neighbourhood context of the textual input. This limits their generalizability. For instance, NMN fail to understand new concepts such as "yellow sphere to the left" even when it is a combination of known concepts from train data: "blue sphere", "yellow cube", and "metallic cube to the left". In this paper, we address this limitation by introducing a language-guided adaptive convolution layer (LG-Conv) into NMN, in which the filter weights of convolutions are explicitly multiplied with a spatially varying language-guided kernel. Our model allows the neural module to adaptively co-attend over potential objects of interest from the visual and textual inputs. Extensive experiments on VQA and REF tasks demonstrate the effectiveness of our approach. Additionally, we propose a new challenging out-of-distribution test split for REF task, which we call C3-Ref+, for explicitly evaluating the NMN's ability to generalize well to adversarial perturbations and unseen combinations of known concepts. Experiments on C3-Ref+ further demonstrate the generalization capabilities of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Grounded Reinforcement Learning for Visual ReasoningGabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain 等NeurIPS 2025 · 被引用 90 次
- EQA-MX: Embodied Question Answering using Multimodal ExpressionMd Mofijul Islam, Alexi Gladstone, Riashat Islam, Tariq IqbalICLR 2024 · 被引用 18 次
- PATRON: Perspective-Aware Multitask Model for Referring Expression Grounding Using Embodied Multimodal CuesMd Mofijul Islam, Alexi Gladstone, Tariq IqbalAAAI 2023 · 被引用 10 次
- In-Context Compositional Generalization for Large Vision-Language ModelsChuanhao Li, Chenchen Jing, Zhen Li, Mingliang Zhai 等EMNLP 2024 · 被引用 1 次
- Exploring the Effect of Primitives for Compositional Generalization in Vision-and-LanguageChuanhao Li, Zhen Li, Chenchen Jing, Yunde Jia 等CVPR 2023
它引用的顶会 Paper7
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 被引用 183 次
- CoCoX: Generating Conceptual and Counterfactual Explanations via Fault-LinesArjun R. Akula, Shuai Wang, Song-Chun ZhuAAAI 2020 · 被引用 102 次
- Closed Loop Neural-Symbolic Learning via Integrating Neural Perception, Grammar Parsing, and Symbolic ReasoningQing Li, Siyuan Huang, Yining Hong, Yixin Chen 等ICML 2020 · 被引用 93 次
- CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA GeneralizationArjun R. Akula, Soravit Changpinyo, Boqing Gong, Piyush Sharma 等EMNLP 2021 · 被引用 18 次
相关 Paper
- Mind the Context: The Impact of Contextualization in Neural Module Networks for Grounding Visual Referring ExpressionsArjun R. Akula, Spandana Gella, Keze Wang, Song-Chun Zhu 等EMNLP 2021 · 被引用 3 次
- How Modular should Neural Module Networks Be for Systematic Generalization?Vanessa D'Amario, Tomotake Sasaki, Xavier BoixNeurIPS 2021 · 被引用 22 次
- TransRefer3D: Entity-and-Relation Aware Transformer for Fine-Grained 3D Visual GroundingDailan He, Yusheng Zhao, Junyu Luo, Tianrui Hui 等ACM MM 2021 · 被引用 81 次
- Cross-Modality Relevance for Reasoning on Language and VisionChen Zheng, Quan Guo, Parisa KordjamshidiACL 2020 · 被引用 33 次
- Language-Guided Visual Aggregation Network for Video Question AnsweringXiao Liang, Di Wang, Quan Wang, Bo Wan 等ACM MM 2023 · 被引用 5 次
