The Neural Race Reduction: Dynamics of Abstraction in Gated Networks
Andrew M. Saxe, Shagun Sodhani, Sam Jay Lewallen
摘要
Our theoretical understanding of deep learning has not kept pace with its empirical success. While network architecture is known to be critical, we do not yet understand its effect on learned representations and network behavior, or how this architecture should reflect task structure.In this work, we begin to address this gap by introducing the Gated Deep Linear Network framework that schematizes how pathways of information flow impact learning dynamics within an architecture. Crucially, because of the gating, these networks can compute nonlinear functions of their input. We derive an exact reduction and, for certain cases, exact solutions to the dynamics of learning. Our analysis demonstrates that the learning dynamics in structured networks can be conceptualized as a neural race with an implicit bias towards shared representations, which then govern the model's ability to systematically generalize, multi-task, and transfer. We validate our key insights on naturalistic datasets and with relaxed assumptions. Taken together, our work gives rise to general hypotheses relating neural architecture to learning and provides a mathematical approach towards understanding the design of more complex architectures and the role of modularity and compositionality in solving real-world problems. The code and results are available at https: //www.saxelab.org/gated-dln .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formationAaditya K. Singh, Ted Moskovitz, Felix Hill, Stephanie C. Y. Chan 等ICML 2024 · 被引用 77 次
- Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learningDaniel Kunin, Allan Raventós, Clémentine C. J. Dominé, Feng Chen 等NeurIPS 2024 · 被引用 48 次
- Parseval Regularization for Continual Reinforcement LearningWesley Chung, Lynn Cherif, Doina Precup, David MegerNeurIPS 2024 · 被引用 27 次
- The Simplicity Bias in Multi-Task RNNs: Shared Attractors, Reuse of Dynamics, and Geometric RepresentationElia Turner, Omri BarakNeurIPS 2023 · 被引用 25 次
- Globally Gated Deep Linear NetworksQianyi Li, Haim SompolinskyNeurIPS 2022 · 被引用 17 次
它引用的顶会 Paper12
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas 等ICML 2020 · 被引用 651 次
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal 等ICML 2021 · 被引用 382 次
- Recurrent Independent MechanismsAnirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani 等ICLR 2021 · 被引用 357 次
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and DepthThao Nguyen, Maithra Raghu, Simon KornblithICLR 2021 · 被引用 323 次
相关 Paper
- Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU NetworksDevon Jarvis, Richard Klein, Benjamin Rosman, Andrew M. SaxeICLR 2025
- Deep Networks Learn Features From Local Discontinuities in the Label FunctionPrithaj Banerjee, Harish Guruprasad Ramaswamy, Mahesh Lorik Yadav, Chandra Shekar LakshminarayananICLR 2025
- Why bigger is not always better: on finite and infinite neural networksLaurence AitchisonICML 2020 · 被引用 59 次
- On The Specialization of Neural ModulesDevon Jarvis, Richard Klein, Benjamin Rosman, Andrew M. SaxeICLR 2023 · 被引用 4 次
- Flexible task abstractions emerge in linear networks with fast and bounded unitsKai Sandbrink, Jan P. Bauer, Alexandra Maria Proca, Andrew M. Saxe 等NeurIPS 2024 · 被引用 7 次
