Is a Modular Architecture Enough?
Sarthak Mittal, Yoshua Bengio, Guillaume Lajoie
Abstract
Inspired from human cognition, machine learning systems are gradually revealing advantages of sparser and more modular architectures. Recent work demonstrates that not only do some modular architectures generalize well, but they also lead to better out-of-distribution generalization, scaling properties, learning speed, and interpretability. A key intuition behind the success of such systems is that the data generating system for most real-world settings is considered to consist of sparsely interacting parts, and endowing models with similar inductive biases will be helpful. However, the field has been lacking in a rigorous quantitative assessment of such systems because these real-world data distributions are complex and unknown. In this work, we provide a thorough assessment of common modular architectures, through the lens of simple and known modular data distributions. We highlight the benefits of modularity and sparsity and reveal insights on the challenges faced while optimizing modular systems. In doing so, we propose evaluation metrics that highlight the benefits of modularity, the regimes in which these benefits are substantial, as well as the sub-optimality of current end-to-end learned modular systems as opposed to their claimed potential. 1 Although a number of recent results hinge on such modular architectures (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1fc38182-a839-436c-bd67-8eb0e0c41eceCited by top-tier papers12
- Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing PolicyPingzhi Li, Zhenyu Zhang, Prateek Yadav, Yi-Lin Sung et al.ICLR 2024 · 97 citations
- Discovering modular solutions that generalize compositionallySimon Schug, Seijin Kobayashi, Yassir Akram, Maciej Wolczyk et al.ICLR 2024 · 24 citations
- Learning Causal Dynamics Models in Object-Oriented EnvironmentsZhongwei Yu, Jingqing Ruan, Dengpeng XingICML 2024 · 4 citations
- On The Specialization of Neural ModulesDevon Jarvis, Richard Klein, Benjamin Rosman, Andrew M. SaxeICLR 2023 · 4 citations
- Self-Supervised Interpretable End-to-End Learning via Latent Functional ModularityHyunki Seong, David Hyunchul ShimICML 2024 · 3 citations
Builds on10
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- A Meta-Transfer Objective for Learning to Disentangle Causal MechanismsYoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke et al.ICLR 2020 · 371 citations
- Recurrent Independent MechanismsAnirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani et al.ICLR 2021 · 357 citations
- Neural Production SystemsAniket Didolkar, Anirudh Goyal, Nan Rosemary Ke, Charles Blundell et al.NeurIPS 2021 · 94 citations
Related papers
- Neural Attentive CircuitsMartin Weiss, Nasim Rahaman, Francesco Locatello, Chris Pal et al.NeurIPS 2022 · 8 citations
- Spatially Structured Recurrent ModulesNasim Rahaman, Anirudh Goyal, Muhammad Waleed Gondal, Manuel Wuthrich et al.ICLR 2021 · 15 citations
- Are Neural Nets Modular? Inspecting Functional Modularity Through Differentiable Weight MasksRóbert Csordás, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2021 · 13 citations
- On the generalization capacity of neural networks during generic multimodal reasoningTakuya Ito, Soham Dan, Mattia Rigotti, James R. Kozloski et al.ICLR 2024 · 4 citations
- Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-LearningTrinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen et al.KDD 2026 · 1 citation
