Lipschitz Lifelong Reinforcement Learning
Erwan Lecarpentier, David Abel, Kavosh Asadi, Yuu Jinnai, Emmanuel Rachelson, Michael L. Littman
Abstract
We consider the problem of knowledge transfer when an agent is facing a series of Reinforcement Learning (RL) tasks. We introduce a novel metric between Markov Decision Processes and establish that close MDPs have close optimal value functions. Formally, the optimal value functions are Lipschitz continuous with respect to the tasks space. These theoretical results lead us to a value-transfer method for Lifelong RL, which we use to build a PAC-MDP algorithm with improved convergence rate. Further, we show the method to experience no negative transfer with high probability. We illustrate the benefits of the method in Lifelong RL experiments. * Kavosh Asadi finished working on this project before joining Amazon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4e1aa98-0a16-46db-a65a-93bb67c790afCited by top-tier papers6
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 244 citations
- Towards Safe Policy Improvement for Non-Stationary MDPsYash Chandak, Scott M. Jordan, Georgios Theocharous, Martha White et al.NeurIPS 2020 · 47 citations
- Resetting the Optimizer in Deep RL: An Empirical StudyKavosh Asadi, Rasool Fakoor, Shoham SabachNeurIPS 2023 · 38 citations
- Policy Caches with Successor FeaturesMark W. Nemecek, Ron ParrICML 2021 · 19 citations
- MINT: Minimal Information Neuro-Symbolic Tree for Objective-Driven Knowledge-Gap Reasoning and Active ElicitationZeyu Fang, Mahdi Imani, Tian LanICML 2026
Related papers
- Performance Bounds for Model and Policy Transfer in Hidden-parameter MDPsHaotian Fu, Jiayu Yao, Omer Gottesman, Finale Doshi-Velez et al.ICLR 2023
- Provably Efficient Lifelong Reinforcement Learning with Linear RepresentationSanae Amani, Lin Yang, Ching-An ChengICLR 2023
- Generalisation in Lifelong Reinforcement Learning through Logical CompositionGeraud Nangue Tasse, Steven James, Benjamin RosmanICLR 2022 · 23 citations
- Lifelong Policy Gradient Learning of Factored Policies for Faster Training Without ForgettingJorge A. Mendez, Boyu Wang, Eric EatonNeurIPS 2020 · 42 citations
- Model-based Lifelong Reinforcement Learning with Bayesian ExplorationHaotian Fu, Shangqun Yu, Michael Littman, George KonidarisNeurIPS 2022 · 17 citations
