Mind the Context: The Impact of Contextualization in Neural Module Networks for Grounding Visual Referring Expressions
Arjun R. Akula, Spandana Gella, Keze Wang, Song-Chun Zhu, Siva Reddy
Abstract
Neural module networks (NMN) are a popular approach for grounding visual referring expressions. Prior implementations of NMN use pre-defined and fixed textual inputs in their module instantiation. This necessitates a large number of modules as they lack the ability to share weights and exploit associations between similar textual contexts (e.g. "dark cube on the left" vs. "black cube on the left"). In this work, we address these limitations and evaluate the impact of contextual clues in improving the performance of NMN models. First, we address the problem of fixed textual inputs by parameterizing the module arguments. This substantially reduce the number of modules in NMN by up to 75% without any loss in performance. Next we propose a method to contextualize our parameterized model to enhance the module's capacity in exploiting the visiolinguistic associations. Our model outperforms the state-of-the-art NMN model on CLEVR-Ref+ dataset with +8.1% improvement in accuracy on the single-referent test set and +4.3% on the full test set. Additionally, we demonstrate that contextualization provides +11.2% and +1.7% improvements in accuracy over prior NMN models on CLO-SURE and NLVR2. We further evaluate the impact of our contextualization by constructing a contrast set for CLEVR-Ref+, which we call CC-Ref+. We significantly outperform the baselines by as much as +10.4% absolute accuracy on CC-Ref+, illustrating the generalization skills of our approach. Our dataset is publicly available at https://github.com/ McGill-NLP/contextual-nmn .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 504ea49a-60cb-407f-91ea-d67ad229434eBuilds on3
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 183 citations
- CoCoX: Generating Conceptual and Counterfactual Explanations via Fault-LinesArjun R. Akula, Shuai Wang, Song-Chun ZhuAAAI 2020 · 102 citations
- REVERIE: Remote Embodied Visual Referring Expression in Real Indoor EnvironmentsYuankai Qi, Qi Wu, Peter Anderson, Xin Wang et al.CVPR 2020
Related papers
- Robust Visual Reasoning via Language Guided Neural Module NetworksArjun R. Akula, Varun Jampani, Soravit Changpinyo, Song-Chun ZhuNeurIPS 2021 · 26 citations
- Graph-Structured Referring Expression Reasoning in the WildSibei Yang, Guanbin Li, Yizhou YuCVPR 2020
- How Modular should Neural Module Networks Be for Systematic Generalization?Vanessa D'Amario, Tomotake Sasaki, Xavier BoixNeurIPS 2021 · 22 citations
- CLEVR-Implicit: A Diagnostic Dataset for Implicit Reasoning in Referring Expression ComprehensionJingwei Zhang, Xin Wu, Yi CaiEMNLP 2023 · 1 citation
- Look Around and Refer: 2D Synthetic Semantics Knowledge Distillation for 3D Visual GroundingEslam Mohamed Bakr, Yasmeen Alsaedy, Mohamed ElhoseinyNeurIPS 2022 · 66 citations
