REPAINT: Knowledge Transfer in Deep Reinforcement Learning
Yunzhe Tao, Sahika Genc, Jonathan Chung, Tao Sun, Sunil Mallya
摘要
Accelerating learning processes for complex tasks by leveraging previously learned tasks has been one of the most challenging problems in reinforcement learning, especially when the similarity between source and target tasks is low. This work proposes REPresentation And INstance Transfer (REPAINT) algorithm for knowledge transfer in deep reinforcement learning. REPAINT not only transfers the representation of a pre-trained teacher policy in the on-policy learning, but also uses an advantage-based experience selection approach to transfer useful samples collected following the teacher policy in the off-policy learning. Our experimental results on several benchmark tasks show that REPAINT significantly reduces the total training time in generic cases of task similarity. In particular, when the source tasks are dissimilar to, or sub-tasks of, the target tasks, REPAINT outperforms other baselines in both training-time reduction and asymptotic performance of return scores.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Cross-Domain Policy Adaptation via Value-Guided Data FilteringKang Xu, Chenjia Bai, Xiaoteng Ma, Dong Wang 等NeurIPS 2023 · 被引用 41 次
- Principled Fast and Meta Knowledge Learners for Continual Reinforcement LearningKe Sun, Hongming Zhang, Jun Jin, Chao Gao 等ICLR 2026 · 被引用 1 次
- Reidentify: Context-Aware Identity Generation for Contextual Multi-Agent Reinforcement LearningZhiwei Xu, Kun Hu, Xin Xin, Weiliang Meng 等ICML 2025
它引用的顶会 Paper1
相关 Paper
- Knowledge Transfer in Multi-Task Deep Reinforcement Learning for Continuous ControlZhiyuan Xu, Kun Wu, Zhengping Che, Jian Tang 等NeurIPS 2020 · 被引用 58 次
- CUP: Critic-Guided Policy ReuseJin Zhang, Siyuan Li, Chongjie ZhangNeurIPS 2022 · 被引用 11 次
- Sharing Knowledge in Multi-Task Deep Reinforcement LearningCarlo D'Eramo, Davide Tateo, Andrea Bonarini, Marcello Restelli 等ICLR 2020 · 被引用 148 次
- Transfer Value Iteration NetworksJunyi Shen, Hankz Hankui Zhuo, Jin Xu, Bin Zhong 等AAAI 2020 · 被引用 7 次
- Composing Task-Agnostic Policies with Deep Reinforcement LearningAhmed Hussain Qureshi, Jacob J. Johnson, Yuzhe Qin, Taylor Henderson 等ICLR 2020 · 被引用 35 次
