Flexible task abstractions emerge in linear networks with fast and bounded units
Kai Sandbrink, Jan P. Bauer, Alexandra Maria Proca, Andrew M. Saxe, Christopher Summerfield, Ali Hummos
摘要
Animals survive in dynamic environments changing at arbitrary timescales, but such data distribution shifts are a challenge to neural networks. To adapt to change, neural systems may change a large number of parameters, which is a slow process involving forgetting past information. In contrast, animals leverage distribution changes to segment their stream of experience into tasks and associate them with internal task abstractions. Animals can then respond flexibly by selecting the appropriate task abstraction. However, how such flexible task abstractions may arise in neural systems remains unknown. Here, we analyze a linear gated network where the weights and gates are jointly optimized via gradient descent, but with neuron-like constraints on the gates including a faster timescale, nonnegativity, and bounded activity. We observe that the weights self-organize into modules specialized for tasks or sub-tasks encountered, while the gates layer forms unique representations that switch the appropriate weight modules (task abstractions). We analytically reduce the learning dynamics to an effective eigenspace, revealing a virtuous cycle: fast adapting gates drive weight specialization by protecting previous knowledge, while weight specialization in turn increases the update rate of the gating layer. Task switching in the gating layer accelerates as a function of curriculum block size and task training, mirroring key findings in cognitive neuroscience. We show that the discovered task abstractions support generalization through both task and subtask composition, and we extend our findings to a non-linear network switching between two tasks. Overall, our work offers a theory of cognitive flexibility in animals as arising from joint gradient descent on synaptic and neural gating in a neural network architecture.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Separating the 'what' and 'how' of compositional computation to enable reuse and continual learningHaozhe Shan, Minni Sun, Lea DunckerNeurIPS 2025 · 被引用 10 次
- Learning dynamics in linear recurrent neural networksAlexandra Maria Proca, Clémentine Carla Juliette Dominé, Murray Shanahan, Pedro A. M. MedianoICML 2025
它引用的顶会 Paper9
- Exact learning dynamics of deep linear networks with prior knowledgeLukas Braun, Clémentine C. J. Dominé, James Fitzgerald, Andrew M. SaxeNeurIPS 2022 · 被引用 75 次
- Contextual Instance Decoupling for Robust Multi-Person Pose EstimationDongkai Wang, Shiliang ZhangCVPR 2022 · 被引用 73 次
- Learning to Combine Top-Down and Bottom-Up Signals in Recurrent Neural Networks with Attention over ModulesSarthak Mittal, Alex Lamb, Anirudh Goyal, Vikram Voleti 等ICML 2020 · 被引用 73 次
- The Neural Race Reduction: Dynamics of Abstraction in Gated NetworksAndrew M. Saxe, Shagun Sodhani, Sam Jay LewallenICML 2022 · 被引用 52 次
- Discovering modular solutions that generalize compositionallySimon Schug, Seijin Kobayashi, Yassir Akram, Maciej Wolczyk 等ICLR 2024 · 被引用 24 次
相关 Paper
- Thalamus: a brain-inspired algorithm for biologically-plausible continual learning and disentangled representationsAli HummosICLR 2023 · 被引用 9 次
- Shaping Sequence Attractor Schema in Recurrent Neural NetworksZhikun Chu, Bo Ho, Xiaolong Zou, Yuanyuan MiNeurIPS 2025
- Artificial Neuronal Ensembles with Learned Context Dependent GatingMatthew J. Tilley, Michelle Miller, David FreedmanICLR 2023 · 被引用 2 次
- A Combinatorial Perspective on Transfer LearningJianan Wang, Eren Sezener, David Budden, Marcus Hutter 等NeurIPS 2020 · 被引用 9 次
- Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU NetworksDevon Jarvis, Richard Klein, Benjamin Rosman, Andrew M. SaxeICLR 2025
