Efficient Multi-task Reinforcement Learning with Cross-Task Policy Guidance
Jinmin He, Kai Li, Yifan Zang, Haobo Fu, Qiang Fu, Junliang Xing, Jian Cheng
Abstract
Multi-task reinforcement learning endeavors to efficiently leverage shared information across various tasks, facilitating the simultaneous learning of multiple tasks. Existing approaches primarily focus on parameter sharing with carefully designed network structures or tailored optimization procedures. However, they overlook a direct and complementary way to exploit cross-task similarities: the control policies of tasks already proficient in some skills can provide explicit guidance for unmastered tasks to accelerate skills acquisition. To this end, we present a novel framework called Cross-Task Policy Guidance (CTPG), which trains a guide policy for each task to select the behavior policy interacting with the environment from all tasks'control policies, generating better training trajectories. In addition, we propose two gating mechanisms to improve the learning efficiency of CTPG: one gate filters out control policies that are not beneficial for guidance, while the other gate blocks tasks that do not necessitate guidance. CTPG is a general framework adaptable to existing parameter sharing approaches. Empirical evaluations demonstrate that incorporating CTPG with these approaches significantly enhances performance in manipulation and locomotion benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d3f0a9f-0f32-470d-b293-0e0cd855e6e6Cited by top-tier papers8
- Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task LearnersMichal Nauman, Marek Cygan, Carmelo Sferrazza, Aviral Kumar et al.NeurIPS 2025 · 26 citations
- Knowledge Diversion for Efficient Morphology Control and Policy TransferFu Feng, Ruixiao Shi, Yucheng Xie, Jianlu Shen et al.ICML 2026 · 1 citation
- TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement LearningHayeong Lee, JunHyeok Oh, Byung-Jun LeeICML 2026
- Words Towards Explainability: Caption Label-Free Learning via Dual Loop Agentic Time Series CaptioningDifei Hou, Jiaqi Yue, Chunhui ZhaoICML 2026
- ARS: Adaptive Reward Scaling for Multi-Task Reinforcement LearningMyungsik Cho, Jongeui Park, Jeonghye Kim, Youngchul SungICML 2025
Builds on11
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- Mastering Complex Control in MOBA Games with Deep Reinforcement LearningDeheng Ye, Zhao Liu, Mingfei Sun, Bei Shi et al.AAAI 2020 · 395 citations
- Multi-Task Reinforcement Learning with Soft ModularizationRuihan Yang, Huazhe Xu, Yi Wu, Xiaolong WangNeurIPS 2020 · 247 citations
- Multi-Task Reinforcement Learning with Context-based RepresentationsShagun Sodhani, Amy Zhang, Joelle PineauICML 2021 · 241 citations
Related papers
- QMP: Q-switch Mixture of Policies for Multi-Task Behavior SharingGrace Zhang, Ayush Jain, Injune Hwang, Shao-Hua Sun et al.ICLR 2025
- Not All Tasks Are Equally Difficult: Multi-Task Deep Reinforcement Learning with Dynamic Depth RoutingJinmin He, Kai Li, Yifan Zang, Haobo Fu et al.AAAI 2024 · 11 citations
- PiCor: Multi-Task Deep Reinforcement Learning with Policy CorrectionFengshuo Bai, Hongming Zhang, Tianyang Tao, Zhiheng Wu et al.AAAI 2023 · 31 citations
- Sample Efficient Myopic Exploration Through Multitask Reinforcement Learning with Diverse TasksZiping Xu, Zifan Xu, Runxuan Jiang, Peter Stone et al.ICLR 2024 · 2 citations
- Minimum Description Length ControlTed Moskovitz, Ta-Chu Kao, Maneesh Sahani, Matt M. BotvinickICLR 2023 · 76 citations
