Robust Knowledge Transfer in Tiered Reinforcement Learning
Jiawei Huang, Niao He
Abstract
In this paper, we study the Tiered Reinforcement Learning setting, a parallel transfer learning framework, where the goal is to transfer knowledge from the low-tier (source) task to the high-tier (target) task to reduce the exploration risk of the latter while solving the two tasks in parallel. Unlike previous work, we do not assume the low-tier and high-tier tasks share the same dynamics or reward functions, and focus on robust knowledge transfer without prior knowledge on the task similarity. We identify a natural and necessary condition called the "Optimal Value Dominance" for our objective. Under this condition, we propose novel online learning algorithms such that, for the high-tier task, it can achieve constant regret on partial states depending on the task similarity and retain near-optimal regret when the two tasks are dissimilar, while for the low-tier task, it can keep near-optimal without making sacrifice. Moreover, we further study the setting with multiple low-tier tasks, and propose a novel transfer source selection mechanism, which can ensemble the information from all low-tier tasks and allow provable benefits on a much larger state-action space. Introduction Comparing with individual learning from scratch, transferring knowledge from other similar tasks or side information has been proven to be an effective way to reduce the exploration risk and improve sample efficiency in Reinforcement Learning (RL). Multi-Task RL (MT-RL) [29] and Transfer RL [26, 18, 37] are two mainstream knowledge transfer frameworks; however, both are subject to limitations when dealing with real-world scenarios. MT-RL studies the setting where a set of similar tasks are solved concurrently, and the main objective is to accelerate the learning by sharing information of all tasks together. However, in practice, in many MT-RL scenarios, the tasks are not equally important and we are more interested in the performance of certain tasks. For example, in robot learning, a few robots are more valuable and hard to fix, while the others are cheaper or just simulators. Most existing works on MT-RL treat all tasks equally and focus primarily on the reduction of the total regret of all tasks as a whole [3, 8, 36, 11] , with no guarantee of improving a particular task. In contrast, transfer RL distinguishes the priority of different tasks by categorizing them into source and target tasks and aims at transferring the knowledge from source tasks (or some side information like value predictors) to facilitate the learning of target tasks [21, 27, 9, 10] . However, a key assumption in transfer RL is that the source task is completely solved before the learning of the target task, and this is not always practical. For example, in some sim-to-real domain, the source task simulator may require a long time to solve [5] , and in some user-interaction scenarios [13], the source and target tasks refer to different user groups and they have to be served simultaneously. In these cases, it's more reasonable to solve the source and target tasks in parallel and transfer the information immediately once available. Recently, [13] proposed a new "parallel knowledge transfer" framework, called Tiered RL, which is promising to fill the gap. Tiered RL considers the case when a source task M Lo and a target task 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a42053a4-c1e3-4660-8ef4-ca42819534a8Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- Bellman-consistent Pessimism for Offline Reinforcement LearningTengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro et al.NeurIPS 2021 · 339 citations
- Multi-Task Reinforcement Learning with Context-based RepresentationsShagun Sodhani, Amy Zhang, Joelle PineauICML 2021 · 241 citations
- Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement LearningTengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong et al.NeurIPS 2021 · 207 citations
- Pessimistic Model-based Offline Reinforcement Learning under Partial CoverageMasatoshi Uehara, Wen SunICLR 2022 · 176 citations
Related papers
- ModelDiff: Symbolic Dynamic Programming for Model-Aware Policy Transfer in Deep Q-LearningXiaotian Liu, Jihwan Jeong, Ayal Taitler, Michael Gimelfarb et al.AAAI 2025
- Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement LearningChi Zhang, Zi-Jia Wang, George K. Atia, Sihong He et al.ICML 2025
- REPAINT: Knowledge Transfer in Deep Reinforcement LearningYunzhe Tao, Sahika Genc, Jonathan Chung, Tao Sun et al.ICML 2021 · 32 citations
- Knowledge Transfer in Multi-Task Deep Reinforcement Learning for Continuous ControlZhiyuan Xu, Kun Wu, Zhengping Che, Jian Tang et al.NeurIPS 2020 · 58 citations
- Transfer Q-Learning with Composite MDP StructuresJinhang Chai, Elynn Y. Chen, Lin YangICML 2025
