On The Specialization of Neural Modules
Devon Jarvis, Richard Klein, Benjamin Rosman, Andrew M. Saxe
Abstract
A number of machine learning models have been proposed with the goal of achieving systematic generalization: the ability to reason about new situations by combining aspects of previous experiences. These models leverage compositional architectures which aim to learn specialized modules dedicated to structures in a task that can be composed to solve novel problems with similar structures. While the compositionality of these architectures is guaranteed by design, the modules specializing is not. Here we theoretically study the ability of network modules to specialize to useful structures in a dataset and achieve systematic generalization. To this end we introduce a minimal space of datasets motivated by practical systematic generalization benchmarks. From this space of datasets we present a mathematical definition of systematicity and study the learning dynamics of linear neural modules when solving components of the task. Our results shed light on the difficulty of module specialization, what is required for modules to successfully specialize, and the necessity of modular architectures to achieve systematicity. Finally, we confirm that the theoretical results in our tractable setting generalize to more complex datasets and non-linear architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d38683bb-6da3-47d7-8388-e46e7b095eb5Cited by top-tier papers10
- Discovering modular solutions that generalize compositionallySimon Schug, Seijin Kobayashi, Yassir Akram, Maciej Wolczyk et al.ICLR 2024 · 24 citations
- Why Do Animals Need Shaping? A Theory of Task Composition and Curriculum LearningJin Hwa Lee, Stefano Sarao Mannelli, Andrew M. SaxeICML 2024 · 15 citations
- Scaling can lead to compositional generalizationFlorian Redhardt, Yassir Akram, Simon SchugNeurIPS 2025 · 11 citations
- Breaking Neural Network Scaling Laws with ModularityAkhilan Boopathy, Sunshine Jiang, William Yue, Jaedong Hwang et al.ICLR 2025 · 1 citation
- Neural Modular Physics for Elastic SimulationYifei Li, Haixu Wu, Zeyi Xu, Tuur Stuyck et al.ICML 2026 · 1 citation
Builds on6
- Recurrent Independent MechanismsAnirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani et al.ICLR 2021 · 357 citations
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt et al.NeurIPS 2020 · 169 citations
- Compositional languages emerge in a neural iterated learning modelYi Ren, Shangmin Guo, Matthieu Labeau, Shay B. Cohen et al.ICLR 2020 · 111 citations
- Is a Modular Architecture Enough?Sarthak Mittal, Yoshua Bengio, Guillaume LajoieNeurIPS 2022 · 62 citations
- Characterizing emergent representations in a space of candidate learning rules for deep networksYinan Cao, Christopher Summerfield, Andrew M. SaxeNeurIPS 2020 · 11 citations
Related papers
- Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated SurveyIvan Vegner, Sydelle de Souza, Valentin Forch, Martha Lewis et al.ACL 2025 · 3 citations
- Iterated learning for emergent systematicity in VQAAnkit Vani, Max Schwarzer, Yuchen Lu, Eeshan Dhekane et al.ICLR 2021 · 27 citations
- Compositional-ARC: Assessing Systematic Generalization in Abstract Spatial ReasoningPhilipp Mondorf, Shijia Zhou, Monica Riedler, Barbara PlankICLR 2026 · 3 citations
- Task-Driven Modular Networks for Zero-Shot Compositional LearningSenthil Purushwalkam, Maximilian Nickel, Abhinav Gupta, Marc'Aurelio RanzatoICCV 2019 · 222 citations
- Break It Down: Evidence for Structural Compositionality in Neural NetworksMichael A. Lepori, Thomas Serre, Ellie PavlickNeurIPS 2023 · 66 citations
