Lifelong Hyper-Policy Optimization with Multiple Importance Sampling Regularization
Pierre Liotet, Francesco Vidaich, Alberto Maria Metelli, Marcello Restelli
Abstract
Learning in a lifelong setting, where the dynamics continually evolve, is a hard challenge for current reinforcement learning algorithms. Yet this would be a much needed feature for practical applications. In this paper, we propose an approach which learns a hyper-policy, whose input is time, that outputs the parameters of the policy to be queried at that time. This hyper-policy is trained to maximize the estimated future performance, efficiently reusing past data by means of importance sampling, at the cost of introducing a controlled bias. We combine the future performance estimate with the past performance to mitigate catastrophic forgetting. To avoid overfitting the collected data, we derive a differentiable variance bound that we embed as a penalization term. Finally, we empirically validate our approach, in comparison with state-of-the-art algorithms, on realistic environments, including water resource management and trading.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Off-Policy Evaluation for Action-Dependent Non-stationary EnvironmentsYash Chandak, Shiv Shankar, Nathaniel D. Bastian, Bruno C. da Silva et al.NeurIPS 2022 · 7 citations
- Truncating Trajectories in Monte Carlo Reinforcement LearningRiccardo Poiani, Alberto Maria Metelli, Marcello RestelliICML 2023 · 6 citations
- Adaptive Instrument Design for Indirect ExperimentsYash Chandak, Shiv Shankar, Vasilis Syrgkanis, Emma BrunskillICLR 2024 · 5 citations
Builds on1
Related papers
- Lifelong Policy Gradient Learning of Factored Policies for Faster Training Without ForgettingJorge A. Mendez, Boyu Wang, Eric EatonNeurIPS 2020 · 42 citations
- Deep Reinforcement Learning amidst Continual Structured Non-StationarityAnnie Xie, James Harrison, Chelsea FinnICML 2021 · 43 citations
- Online Reinforcement Learning in Non-Stationary Context-Driven EnvironmentsPouya Hamadanian, Arash Nasr-Esfahany, Malte Schwarzkopf, Siddhartha Sen et al.ICLR 2025
- Pausing Policy Learning in Non-stationary Reinforcement LearningHyunin Lee, Ming Jin, Javad Lavaei, Somayeh SojoudiICML 2024 · 4 citations
- DRAE: Dynamic Retrieval-Augmented Expert Networks for Lifelong Learning and Task Adaptation in RoboticsYayu Long, Kewei Chen, Long Jin, Mingsheng ShangACL 2025 · 6 citations
