Not All Tasks Are Equally Difficult: Multi-Task Deep Reinforcement Learning with Dynamic Depth Routing
Jinmin He, Kai Li, Yifan Zang, Haobo Fu, Qiang Fu, Junliang Xing, Jian Cheng
Abstract
Multi-task reinforcement learning endeavors to accomplish a set of different tasks with a single policy. To enhance data efficiency by sharing parameters across multiple tasks, a common practice segments the network into distinct modules and trains a routing network to recombine these modules into task-specific policies. However, existing routing approaches employ a fixed number of modules for all tasks, neglecting that tasks with varying difficulties commonly require varying amounts of knowledge. This work presents a Dynamic Depth Routing (D2R) framework, which learns strategic skipping of certain intermediate modules, thereby flexibly choosing different numbers of modules for each task. Under this framework, we further introduce a ResRouting method to address the issue of disparate routing paths between behavior and target policies during off-policy training. In addition, we design an automatic route-balancing mechanism to encourage continued routing exploration for unmastered tasks without disturbing the routing of mastered ones. We conduct extensive experiments on various robotics manipulation tasks in the Meta-World benchmark, where D2R achieves state-of-the-art performance with significantly improved learning efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cae69fd-463b-4012-8f75-85bac13af7ecCited by top-tier papers9
- Learning and Planning Multi-Agent Tasks via an MoE-based World ModelZijie Zhao, Zhongyue Zhao, Kaixuan Xu, Yuqian Fu et al.NeurIPS 2025 · 12 citations
- Efficient Multi-task Reinforcement Learning with Cross-Task Policy GuidanceJinmin He, Kai Li, Yifan Zang, Haobo Fu et al.NeurIPS 2024 · 11 citations
- Centralized Reward Agent for Knowledge Sharing and Transfer in Multi-Task Reinforcement LearningHaozhe Ma, Zhengding Luo, Thanh Vinh Vo, Kuankuan Sima et al.NeurIPS 2025 · 9 citations
- Soft Conflict-Resolution Decision Transformer for Offline Multi-Task Reinforcement LearningShudong Wang, Xinfei Wang, Chenhao Zhang, Shanchen Pang et al.AAAI 2026 · 2 citations
- Mixture-of-World Models: Scaling Multi-Task Reinforcement Learning with Modular Latent DynamicsBoxuan Zhang, Weipu Zhang, Zhaohan Feng, Wei Xiao et al.ICLR 2026 · 1 citation
Builds on8
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- Mastering Complex Control in MOBA Games with Deep Reinforcement LearningDeheng Ye, Zhao Liu, Mingfei Sun, Bei Shi et al.AAAI 2020 · 395 citations
- Multi-Task Reinforcement Learning with Soft ModularizationRuihan Yang, Huazhe Xu, Yi Wu, Xiaolong WangNeurIPS 2020 · 247 citations
Related papers
- PiCor: Multi-Task Deep Reinforcement Learning with Policy CorrectionFengshuo Bai, Hongming Zhang, Tianyang Tao, Zhiheng Wu et al.AAAI 2023 · 31 citations
- Hard Tasks First: Multi-Task Reinforcement Learning Through Task SchedulingMyungsik Cho, Jongeui Park, Suyoung Lee, Youngchul SungICML 2024 · 4 citations
- QMP: Q-switch Mixture of Policies for Multi-Task Behavior SharingGrace Zhang, Ayush Jain, Injune Hwang, Shao-Hua Sun et al.ICLR 2025
- AdaShare: Learning What To Share For Efficient Deep Multi-Task LearningXimeng Sun, Rameswar Panda, Rogério Feris, Kate SaenkoNeurIPS 2020 · 337 citations
- ARS: Adaptive Reward Scaling for Multi-Task Reinforcement LearningMyungsik Cho, Jongeui Park, Jeonghye Kim, Youngchul SungICML 2025
