How Modular should Neural Module Networks Be for Systematic Generalization?
Vanessa D'Amario, Tomotake Sasaki, Xavier Boix
Abstract
Neural Module Networks (NMNs) aim at Visual Question Answering (VQA) via composition of modules that tackle a sub-task. NMNs are a promising strategy to achieve systematic generalization, ie. overcoming biasing factors in the training distribution. However, the aspects of NMNs that facilitate systematic generalization are not fully understood. In this paper, we demonstrate that the degree of modularity of the NMN have large influence on systematic generalization. In a series of experiments on three VQA datasets (VQA-MNIST, SQOOP, and CLEVR-CoGenT), our results reveal that tuning the degree of modularity, especially at the image encoder stage, reaches substantially higher systematic generalization. These findings lead to new NMN architectures that outperform previous ones in terms of systematic generalization. IMAGE ENCODER(S) to obtain visual features INTERMEDIATE MODULE(S) to carry out sub-tasks CLASSIFIER(S) to provide an answer group 1 group 1 [arg] group 1 group | group | group sub-task sub-task sub-task sub-task | sub-task | sub-task all | group | group all group 1 [arg] group 1 all | all | all all all [arg] all group 1 group | all | all all [arg] all '2' left_of green '2' all [green] all all ['2'] all [left_of] color category Question: "Is the green object left of '2'?" Program layout: group | group | group sub-task | sub-task | sub-task all | group | group all | all | all group | all | all left_of green '2'
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21e68003-21f8-461e-b159-a6099f9c223eCited by top-tier papers6
- Learning to Compose: Improving Object Centric Learning by Injecting CompositionalityWhie Jung, Jaehoon Yoo, Sungjin Ahn, Seunghoon HongICLR 2024 · 10 citations
- Object-Category Aware Reinforcement LearningQi Yi, Rui Zhang, Shaohui Peng, Jiaming Guo et al.NeurIPS 2022 · 7 citations
- Improved Object-Centric Diffusion Learning with Registers and Contrastive AlignmentBac Nguyen, Yuhta Takida, Naoki Murata, Chieh-Hsin Lai et al.ICLR 2026 · 3 citations
- DNN Modularization via Activation-Driven TrainingTuan Ngo, Abid Hassan, Saad Shafiq, Nenad MedvidovićICSE 2026 · 2 citations
- Breaking Neural Network Scaling Laws with ModularityAkhilan Boopathy, Sunshine Jiang, William Yue, Jaedong Hwang et al.ICLR 2025 · 1 citation
Builds on3
- Task-Driven Modular Networks for Zero-Shot Compositional LearningSenthil Purushwalkam, Maximilian Nickel, Abhinav Gupta, Marc'Aurelio RanzatoICCV 2019 · 222 citations
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt et al.NeurIPS 2020 · 169 citations
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha et al.NeurIPS 2020 · 163 citations
Related papers
- Robust Visual Reasoning via Language Guided Neural Module NetworksArjun R. Akula, Varun Jampani, Soravit Changpinyo, Song-Chun ZhuNeurIPS 2021 · 26 citations
- Neural Module Networks for Reasoning over TextNitish Gupta, Kevin Lin, Dan Roth, Sameer Singh et al.ICLR 2020 · 134 citations
- Toward Multi-Granularity Decision-Making: Explicit Visual Reasoning with Hierarchical KnowledgeYifeng Zhang, Shi Chen, Qi ZhaoICCV 2023 · 6 citations
- Multimodal Neural Graph Memory Networks for Visual Question AnsweringMahmoud KhademiACL 2020 · 35 citations
- Mind the Context: The Impact of Contextualization in Neural Module Networks for Grounding Visual Referring ExpressionsArjun R. Akula, Spandana Gella, Keze Wang, Song-Chun Zhu et al.EMNLP 2021 · 3 citations
