Discovering modular solutions that generalize compositionally
Simon Schug, Seijin Kobayashi, Yassir Akram, Maciej Wolczyk, Alexandra Maria Proca, Johannes von Oswald, Razvan Pascanu, João Sacramento, Angelika Steger
摘要
Many complex tasks can be decomposed into simpler, independent parts. Discovering such underlying compositional structure has the potential to enable compositional generalization. Despite progress, our most powerful systems struggle to compose flexibly. It therefore seems natural to make models more modular to help capture the compositional nature of many tasks. However, it is unclear under which circumstances modular systems can discover hidden compositional structure. To shed light on this question, we study a teacher-student setting with a modular teacher where we have full control over the composition of ground truth modules. This allows us to relate the problem of compositional generalization to that of identification of the underlying modules. In particular we study modularity in hypernetworks representing a general class of multiplicative interactions. We show theoretically that identification up to linear transformation purely from demonstrations is possible without having to learn an exponential number of module combinations. We further demonstrate empirically that under the theoretically identified conditions, meta-learning from finite data can discover modular policies that generalize compositionally in a number of complex environments. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Why Do Animals Need Shaping? A Theory of Task Composition and Curriculum LearningJin Hwa Lee, Stefano Sarao Mannelli, Andrew M. SaxeICML 2024 · 被引用 15 次
- Scaling can lead to compositional generalizationFlorian Redhardt, Yassir Akram, Simon SchugNeurIPS 2025 · 被引用 11 次
- Flexible task abstractions emerge in linear networks with fast and bounded unitsKai Sandbrink, Jan P. Bauer, Alexandra Maria Proca, Andrew M. Saxe 等NeurIPS 2024 · 被引用 7 次
- Composing Linear Layers from IrreduciblesTravis Pence, Daisuke Yamada, Vikas SinghNeurIPS 2025 · 被引用 1 次
- Compositional Risk MinimizationDivyat Mahajan, Mohammad Pezeshki, Charles Arnal, Ioannis Mitliagkas 等ICML 2025
它引用的顶会 Paper26
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAMLAniruddh Raghu, Maithra Raghu, Samy Bengio, Oriol VinyalsICLR 2020 · 被引用 736 次
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman 等ICLR 2020 · 被引用 401 次
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei 等ICLR 2023 · 被引用 318 次
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya 等NeurIPS 2023 · 被引用 295 次
相关 Paper
- Breaking Neural Network Scaling Laws with ModularityAkhilan Boopathy, Sunshine Jiang, William Yue, Jaedong Hwang 等ICLR 2025 · 被引用 1 次
- On The Specialization of Neural ModulesDevon Jarvis, Richard Klein, Benjamin Rosman, Andrew M. SaxeICLR 2023 · 被引用 4 次
- Attention as a HypernetworkSimon Schug, Seijin Kobayashi, Yassir Akram, João Sacramento 等ICLR 2025
- Learning by Analogy: A Causal Framework for Compositional GeneralizationLingjing Kong, Shaoan Xie, Yang Jiao, Yetian Chen 等CVPR 2026
- Break It Down: Evidence for Structural Compositionality in Neural NetworksMichael A. Lepori, Thomas Serre, Ellie PavlickNeurIPS 2023 · 被引用 66 次
