Continual Task Allocation in Meta-Policy Network via Sparse Prompting
Yijun Yang, Tianyi Zhou, Jing Jiang, Guodong Long, Yuhui Shi
Abstract
How to train a generalizable meta-policy by continually learning a sequence of tasks? It is a natural human skill yet challenging to achieve by current reinforcement learning: the agent is expected to quickly adapt to new tasks (plasticity) meanwhile retaining the common knowledge from previous tasks (stability). We address it by "Continual Task Allocation via Sparse Prompting (CoTASP)", which learns over-complete dictionaries to produce sparse masks as prompts extracting a sub-network for each task from a meta-policy network. CoTASP trains a policy for each task by optimizing the prompts and the subnetwork weights alternatively. The dictionary is then updated to align the optimized prompts with tasks' embedding, thereby capturing tasks' semantic correlations. Hence, relevant tasks share more neurons in the meta-policy network due to similar prompts while cross-task interference causing forgetting is effectively restrained. Given a metapolicy and dictionaries trained on previous tasks, new task adaptation reduces to highly efficient sparse prompting and sub-network finetuning. In experiments, CoTASP achieves a promising plasticity-stability trade-off without storing or replaying any past tasks' experiences. It outperforms existing continual and multi-task RL methods on all seen tasks, forgetting reduction, and generalization to unseen tasks. Our code is available at https://github.com/stevenyangyj/CoTASP
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e79caff8-4b56-43fa-ade1-665920d11ab9Cited by top-tier papers8
- Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorldYijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao et al.CVPR 2024 · 23 citations
- Continual Knowledge Adaptation for Reinforcement LearningJinwu Hu, Zihao Lian, Zhiquan Wen, Chenghao Li et al.NeurIPS 2025 · 8 citations
- Tackling Continual Offline RL through Selective Weights Activation on Aligned SpacesJifeng Hu, Sili Huang, Li Shen, Zhejian Yang et al.NeurIPS 2025 · 2 citations
- Principled Fast and Meta Knowledge Learners for Continual Reinforcement LearningKe Sun, Hongming Zhang, Jun Jin, Chao Gao et al.ICLR 2026 · 1 citation
- Continual Reinforcement Learning by Planning with Online World ModelsZichen Liu, Guoji Fu, Chao Du, Wee Sun Lee et al.ICML 2025
Builds on17
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 569 citations
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
- Achieving Forgetting Prevention and Knowledge Transfer in Continual LearningZixuan Ke, Bing Liu, Nianzu Ma, Hu Xu et al.NeurIPS 2021 · 167 citations
- Forget-free Continual Learning with Winning SubnetworksHaeyong Kang, Rusty John Lloyd Mina, Sultan Rizky Hikmawan Madjid, Jaehong Yoon et al.ICML 2022 · 159 citations
Related papers
- CoMPS: Continual Meta Policy SearchGlen Berseth, Zhiwei Zhang, Grace Zhang, Chelsea Finn et al.ICLR 2022 · 19 citations
- Building a Subspace of Policies for Scalable Continual LearningJean-Baptiste Gaya, Thang Doan, Lucas Caccia, Laure Soulier et al.ICLR 2023 · 3 citations
- Mixture of Meta-Policies for Cross-Environment Meta-Reinforcement LearningXinyu Liu, Qingyu Zeng, Chenwei Tang, Jiancheng LvKDD 2026
- Learning to Modulate pre-trained Models in RLThomas Schmied, Markus Hofmarcher, Fabian Paischer, Razvan Pascanu et al.NeurIPS 2023 · 34 citations
- Parameter-efficient Continual Learning for Enhancing Plasticity without Forgetting under Limited Model CapacityYitian Chen, Shigeng Zhang, Xuan Liu, Mingming Lu et al.CVPR 2026
