Lifelong Policy Gradient Learning of Factored Policies for Faster Training Without Forgetting
Jorge A. Mendez, Boyu Wang, Eric Eaton
Abstract
Policy gradient methods have shown success in learning control policies for high-dimensional dynamical systems. Their biggest downside is the amount of exploration they require before yielding high-performing policies. In a lifelong learning setting, in which an agent is faced with multiple consecutive tasks over its lifetime, reusing information from previously seen tasks can substantially accelerate the learning of new tasks. We provide a novel method for lifelong policy gradient learning that trains lifelong function approximators directly via policy gradients, allowing the agent to benefit from accumulated knowledge throughout the entire training process. We show empirically that our algorithm learns faster and converges to better policies than single-task and lifelong learning baselines, and completely avoids catastrophic forgetting on a variety of challenging domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 320a6008-c9f0-4bbc-8d4e-242e9d86d9eeCited by top-tier papers9
- Continual World: A Robotic Benchmark For Continual Reinforcement LearningMaciej Wolczyk, Michal Zajac, Razvan Pascanu, Lukasz Kucinski et al.NeurIPS 2021 · 152 citations
- Same State, Different Task: Continual Reinforcement Learning without InterferenceSamuel Kessler, Jack Parker-Holder, Philip J. Ball, Stefan Zohren et al.AAAI 2022 · 57 citations
- Modular Lifelong Reinforcement Learning via Neural CompositionJorge A. Mendez, Harm van Seijen, Eric EatonICLR 2022 · 51 citations
- Autonomous Reinforcement Learning: Formalism and BenchmarkingArchit Sharma, Kelvin Xu, Nikhil Sardana, Abhishek Gupta et al.ICLR 2022 · 39 citations
- Fast TRAC: A Parameter-Free Optimizer for Lifelong Reinforcement LearningAneesh Muppidi, Zhiyu Zhang, Heng YangNeurIPS 2024 · 19 citations
Builds on2
- Uncertainty-guided Continual Learning with Bayesian Neural NetworksSayna Ebrahimi, Mohamed Elhoseiny, Trevor Darrell, Marcus RohrbachICLR 2020 · 211 citations
- Functional Regularisation for Continual Learning with Gaussian ProcessesMichalis K. Titsias, Jonathan Schwarz, Alexander G. de G. Matthews, Razvan Pascanu et al.ICLR 2020 · 209 citations
Related papers
- Lifelong Hyper-Policy Optimization with Multiple Importance Sampling RegularizationPierre Liotet, Francesco Vidaich, Alberto Maria Metelli, Marcello RestelliAAAI 2022 · 10 citations
- Deep Reinforcement Learning amidst Continual Structured Non-StationarityAnnie Xie, James Harrison, Chelsea FinnICML 2021 · 43 citations
- Provably Efficient Lifelong Reinforcement Learning with Linear RepresentationSanae Amani, Lin Yang, Ching-An ChengICLR 2023
- A Boolean Task Algebra for Reinforcement LearningGeraud Nangue Tasse, Steven James, Benjamin RosmanNeurIPS 2020 · 71 citations
- Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot ManipulationYuanqi Yao, Siao Liu, Haoming Song, Delin Qu et al.CVPR 2025
