PiCor: Multi-Task Deep Reinforcement Learning with Policy Correction
Fengshuo Bai, Hongming Zhang, Tianyang Tao, Zhiheng Wu, Yanna Wang, Bo Xu
摘要
Multi-task deep reinforcement learning (DRL) ambitiously aims to train a general agent that masters multiple tasks simultaneously. However, varying learning speeds of different tasks compounding with negative gradient interference makes policy learning inefficient. In this work, we propose PiCor, an efficient multi-task DRL framework that splits learning into policy optimization and policy correction phases. The policy optimization phase improves the policy by any DRL algothrim on the sampled single task without considering other tasks. The policy correction phase first constructs a performance constraint set with adaptive weight adjusting. Then the intermediate policy learned by the first phase is constrained to the set, which controls the negative interference and balances the learning speeds across tasks. Empirically, we demonstrate that PiCor outperforms previous methods and significantly improves sample efficiency on simulated robotic manipulation and continuous control tasks. We additionally show that adaptive weight adjusting can further improve data efficiency and performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)Zhenjie Yang, Xiaosong Jia, Qifeng Li, Xue Yang 等NeurIPS 2025 · 被引用 65 次
- Reinforcing LLM Agents via Policy Optimization with Action DecompositionMuning Wen, Ziyu Wan, Jun Wang, Weinan Zhang 等NeurIPS 2024 · 被引用 31 次
- RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted BehaviorsFengshuo Bai, Runze Liu, Yali Du, Ying Wen 等AAAI 2025 · 被引用 15 次
- PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic ManipulationRunze Liu, Yali Du, Fengshuo Bai, Jiafei Lyu 等ICML 2024 · 被引用 10 次
- Centralized Reward Agent for Knowledge Sharing and Transfer in Multi-Task Reinforcement LearningHaozhe Ma, Zhengding Luo, Thanh Vinh Vo, Kuankuan Sima 等NeurIPS 2025 · 被引用 9 次
它引用的顶会 Paper4
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Multi-Task Reinforcement Learning with Soft ModularizationRuihan Yang, Huazhe Xu, Yi Wu, Xiaolong WangNeurIPS 2020 · 被引用 247 次
- Sharing Knowledge in Multi-Task Deep Reinforcement LearningCarlo D'Eramo, Davide Tateo, Andrea Bonarini, Marcello Restelli 等ICLR 2020 · 被引用 148 次
- Knowledge Transfer in Multi-Task Deep Reinforcement Learning for Continuous ControlZhiyuan Xu, Kun Wu, Zhengping Che, Jian Tang 等NeurIPS 2020 · 被引用 58 次
相关 Paper
- HyMTRL: A Hybrid Multi-Task Reinforcement Learning Framework via Phased Policy EvolutionJinmin He, Kai Li, Xiaoyi Dong, Yifan Zang 等ICML 2026
- Not All Tasks Are Equally Difficult: Multi-Task Deep Reinforcement Learning with Dynamic Depth RoutingJinmin He, Kai Li, Yifan Zang, Haobo Fu 等AAAI 2024 · 被引用 11 次
- Efficient Multi-task Reinforcement Learning with Cross-Task Policy GuidanceJinmin He, Kai Li, Yifan Zang, Haobo Fu 等NeurIPS 2024 · 被引用 11 次
- FAMO: Fast Adaptive Multitask OptimizationBo Liu, Yihao Feng, Peter Stone, Qiang LiuNeurIPS 2023 · 被引用 127 次
- Scalable Multi-Objective and Meta Reinforcement Learning via Gradient EstimationZhenshuo Zhang, Minxuan Duan, Youran Ye, Hongyang R. ZhangAAAI 2026 · 被引用 3 次
