How Modular should Neural Module Networks Be for Systematic Generalization?
Vanessa D'Amario, Tomotake Sasaki, Xavier Boix
摘要
Neural Module Networks (NMNs) aim at Visual Question Answering (VQA) via composition of modules that tackle a sub-task. NMNs are a promising strategy to achieve systematic generalization, ie. overcoming biasing factors in the training distribution. However, the aspects of NMNs that facilitate systematic generalization are not fully understood. In this paper, we demonstrate that the degree of modularity of the NMN have large influence on systematic generalization. In a series of experiments on three VQA datasets (VQA-MNIST, SQOOP, and CLEVR-CoGenT), our results reveal that tuning the degree of modularity, especially at the image encoder stage, reaches substantially higher systematic generalization. These findings lead to new NMN architectures that outperform previous ones in terms of systematic generalization. IMAGE ENCODER(S) to obtain visual features INTERMEDIATE MODULE(S) to carry out sub-tasks CLASSIFIER(S) to provide an answer group 1 group 1 [arg] group 1 group | group | group sub-task sub-task sub-task sub-task | sub-task | sub-task all | group | group all group 1 [arg] group 1 all | all | all all all [arg] all group 1 group | all | all all [arg] all '2' left_of green '2' all [green] all all ['2'] all [left_of] color category Question: "Is the green object left of '2'?" Program layout: group | group | group sub-task | sub-task | sub-task all | group | group all | all | all group | all | all left_of green '2'
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Learning to Compose: Improving Object Centric Learning by Injecting CompositionalityWhie Jung, Jaehoon Yoo, Sungjin Ahn, Seunghoon HongICLR 2024 · 被引用 10 次
- Object-Category Aware Reinforcement LearningQi Yi, Rui Zhang, Shaohui Peng, Jiaming Guo 等NeurIPS 2022 · 被引用 7 次
- Improved Object-Centric Diffusion Learning with Registers and Contrastive AlignmentBac Nguyen, Yuhta Takida, Naoki Murata, Chieh-Hsin Lai 等ICLR 2026 · 被引用 3 次
- DNN Modularization via Activation-Driven TrainingTuan Ngo, Abid Hassan, Saad Shafiq, Nenad MedvidovićICSE 2026 · 被引用 2 次
- Breaking Neural Network Scaling Laws with ModularityAkhilan Boopathy, Sunshine Jiang, William Yue, Jaedong Hwang 等ICLR 2025 · 被引用 1 次
它引用的顶会 Paper3
- Task-Driven Modular Networks for Zero-Shot Compositional LearningSenthil Purushwalkam, Maximilian Nickel, Abhinav Gupta, Marc'Aurelio RanzatoICCV 2019 · 被引用 222 次
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt 等NeurIPS 2020 · 被引用 169 次
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha 等NeurIPS 2020 · 被引用 163 次
相关 Paper
- Robust Visual Reasoning via Language Guided Neural Module NetworksArjun R. Akula, Varun Jampani, Soravit Changpinyo, Song-Chun ZhuNeurIPS 2021 · 被引用 26 次
- Neural Module Networks for Reasoning over TextNitish Gupta, Kevin Lin, Dan Roth, Sameer Singh 等ICLR 2020 · 被引用 134 次
- Toward Multi-Granularity Decision-Making: Explicit Visual Reasoning with Hierarchical KnowledgeYifeng Zhang, Shi Chen, Qi ZhaoICCV 2023 · 被引用 6 次
- Multimodal Neural Graph Memory Networks for Visual Question AnsweringMahmoud KhademiACL 2020 · 被引用 35 次
- Mind the Context: The Impact of Contextualization in Neural Module Networks for Grounding Visual Referring ExpressionsArjun R. Akula, Spandana Gella, Keze Wang, Song-Chun Zhu 等EMNLP 2021 · 被引用 3 次
