ModelDiff: Symbolic Dynamic Programming for Model-Aware Policy Transfer in Deep Q-Learning
Xiaotian Liu, Jihwan Jeong, Ayal Taitler, Michael Gimelfarb, Scott Sanner
Abstract
Despite significant recent advances in the field of Deep Reinforcement Learning (DRL), such methods typically incur high cost of training to learn effective policies, thus posing cost and safety challenges in many practical applications. To improve the learning efficiency of (D)RL methods, transfer learning (TL) has emerged as a promising approach to leverage prior experience on a source domain to speed learning on a new, but related, target domain. In this paper, we take a novel model-informed approach to TL in DRL by assuming that we have knowledge of both the source and target domain models (which would be the case in the prevalent setting of DRL with simulators). While directly solving either the source or target MDP via solution methods like value iteration is computationally prohibitive, we exploit the fact that if the target and source MDPs differ only due to a small structural change in their rewards, we can apply structured value iteration methods in a procedure we term ModelDiff to solve the much smaller target-source ``Diff'' MDP for a reasonable horizon. This ModelDiff approach can then be integrated into extensions of standard DRL algorithms like ModelDiff (MD) DQN, where it provides enhanced provable lower bound guidance to DQN that often speeds convergence for the positive transfer case while critically avoiding decelerated learning in the negative transfer case. Experiments show that MD-DQN matches or outperforms existing TL methods and baselines in both positive and negative transfer settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- Robust Knowledge Transfer in Tiered Reinforcement LearningJiawei Huang, Niao HeNeurIPS 2023 · 1 citation
- Transfer Learning for Efficient Iterative Safety ValidationAnthony Corso, Mykel J. KochenderferAAAI 2021 · 6 citations
- Transfer Value Iteration NetworksJunyi Shen, Hankz Hankui Zhuo, Jin Xu, Bin Zhong et al.AAAI 2020 · 7 citations
- Evolving Reinforcement Learning AlgorithmsJohn D. Co-Reyes, Yingjie Miao, Daiyi Peng, Esteban Real et al.ICLR 2021 · 19 citations
- An Efficient Transfer Learning Framework for Multiagent Reinforcement LearningTianpei Yang, Weixun Wang, Hongyao Tang, Jianye Hao et al.NeurIPS 2021 · 34 citations
