Why Do Animals Need Shaping? A Theory of Task Composition and Curriculum Learning
Jin Hwa Lee, Stefano Sarao Mannelli, Andrew M. Saxe
Abstract
Diverse studies in systems neuroscience begin with extended periods of curriculum training known as `shaping' procedures. These involve progressively studying component parts of more complex tasks, and can make the difference between learning a task quickly, slowly or not at all. Despite the importance of shaping to the acquisition of complex tasks, there is as yet no theory that can help guide the design of shaping procedures, or more fundamentally, provide insight into its key role in learning. Modern deep reinforcement learning systems might implicitly learn compositional primitives within their multilayer policy networks. Inspired by these models, we propose and analyse a model of deep policy gradient learning of simple compositional reinforcement learning tasks. Using the tools of statistical physics, we solve for exact learning dynamics and characterise different learning strategies including primitives pre-training, in which task primitives are studied individually before learning compositional tasks. We find a complex interplay between task complexity and the efficacy of shaping strategies. Overall, our theory provides an analytical understanding of the benefits of shaping in a class of compositional tasks and a quantitative account of how training protocols can disclose useful task primitives, ultimately yielding faster and more robust learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ece1385-844a-4a11-aa03-206d1a3be867Cited by top-tier papers8
- Scaling can lead to compositional generalizationFlorian Redhardt, Yassir Akram, Simon SchugNeurIPS 2025 · 11 citations
- Flexible task abstractions emerge in linear networks with fast and bounded unitsKai Sandbrink, Jan P. Bauer, Alexandra Maria Proca, Andrew M. Saxe et al.NeurIPS 2024 · 7 citations
- Influence Dynamics and Stagewise Data AttributionJin Hwa Lee, Matthew Smith, Maxwell Adam, Jesse HooglandICLR 2026 · 5 citations
- Learning to Solve Complex Problems via Dataset DecompositionWanru Zhao, Lucas Page-Caccia, Zhengyan Shi, Minseon Kim et al.NeurIPS 2025 · 3 citations
- Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU NetworksDevon Jarvis, Richard Klein, Benjamin Rosman, Andrew M. SaxeICLR 2025
Builds on15
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu et al.ICLR 2020 · 751 citations
- Recurrent Independent MechanismsAnirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani et al.ICLR 2021 · 357 citations
- When Do Curricula Work?Xiaoxia Wu, Ethan Dyer, Behnam NeyshaburICLR 2021 · 141 citations
- Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic TaskMaya Okawa, Ekdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2023 · 113 citations
- Continual Learning in the Teacher-Student Setup: Impact of Task SimilaritySebastian Lee, Sebastian Goldt, Andrew M. SaxeICML 2021 · 98 citations
Related papers
- Utilizing Prior Solutions for Reward Shaping and Composition in Entropy-Regularized Reinforcement LearningJacob Adamczyk, Argenis Arriojas, Stas Tiomkin, Rahul V. KulkarniAAAI 2023 · 13 citations
- Curriculum learning as a tool to uncover learning principles in the brainDaniel R. Kepple, Rainer Engelken, Kanaka RajanICLR 2022 · 21 citations
- Compositional generalization through abstract representations in human and artificial neural networksTakuya Ito, Tim Klinger, Douglas Schultz, John Murray et al.NeurIPS 2022 · 65 citations
- Programmatic Reinforcement Learning without OraclesWenjie Qiu, He ZhuICLR 2022 · 42 citations
- Shaping Sequence Attractor Schema in Recurrent Neural NetworksZhikun Chu, Bo Ho, Xiaolong Zou, Yuanyuan MiNeurIPS 2025
