Learning a subspace of policies for online adaptation in Reinforcement Learning
Jean-Baptiste Gaya, Laure Soulier, Ludovic Denoyer
Abstract
Deep Reinforcement Learning (RL) is mainly studied in a setting where the training and the testing environments are similar. But in many practical applications, these environments may differ. For instance, in control systems, the robot(s) on which a policy is learned might differ from the robot(s) on which a policy will run. It can be caused by different internal factors (e.g., calibration issues, system attrition, defective modules) or also by external changes (e.g., weather conditions). There is a need to develop RL methods that generalize well to variations of the training conditions. In this article, we consider the simplest yet hard to tackle generalization setting where the test environment is unknown at train time, forcing the agent to adapt to the system's new dynamics. This online adaptation process can be computationally expensive (e.g., fine-tuning) and cannot rely on meta-RL techniques since there is just a single train environment. To do so, we propose an approach where we learn a subspace of policies within the parameter space. This subspace contains an infinite number of policies that are trained to solve the training environment while having different parameter values. As a consequence, two policies in that subspace process information differently and exhibit different behaviors when facing variations of the train environment. Our experiments 1 carried out over a large variety of benchmarks compare our approach with baselines, including diversity-based methods. In comparison, our approach is simple to tune, does not need any extra component (e.g., discriminator) and learns policies able to gather a high reward on unseen environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d9f4031-e6dd-4d13-b18e-88cdccfc7ac6Cited by top-tier papers7
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
- WARM: On the Benefits of Weight Averaged Reward ModelsAlexandre Ramé, Nino Vieillard, Léonard Hussenot, Robert Dadashi et al.ICML 2024 · 145 citations
- Identifying Policy Gradient SubspacesJan Schneider, Pierre Schumacher, Simon Guist, Le Chen et al.ICLR 2024 · 7 citations
- Neuroevolution is a Competitive Alternative to Reinforcement Learning for Skill DiscoveryFélix Chalumeau, Raphaël Boige, Bryan Lim, Valentin Macé et al.ICLR 2023 · 6 citations
- Rewiring Neurons in Non-Stationary EnvironmentsZhicheng Sun, Yadong MuNeurIPS 2023 · 4 citations
Builds on7
- Robust Reinforcement Learning on State Observations with Learned Optimal AdversaryHuan Zhang, Hongge Chen, Duane S. Boning, Cho-Jui HsiehICLR 2021 · 212 citations
- Linear Mode Connectivity in Multitask and Continual LearningSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Dilan Görür, Razvan Pascanu et al.ICLR 2021 · 176 citations
- Robust Deep Reinforcement Learning through Adversarial LossTuomas P. Oikarinen, Wang Zhang, Alexandre Megretski, Luca Daniel et al.NeurIPS 2021 · 134 citations
- One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RLSaurabh Kumar, Aviral Kumar, Sergey Levine, Chelsea FinnNeurIPS 2020 · 109 citations
- Learning Neural Network SubspacesMitchell Wortsman, Maxwell Horton, Carlos Guestrin, Ali Farhadi et al.ICML 2021 · 101 citations
Related papers
- Improving Generalization in Meta-RL with Imaginary Tasks from Latent Dynamics MixtureSuyoung Lee, Sae-Young ChungNeurIPS 2021 · 23 citations
- Monotonic Robust Policy Optimization with Model DiscrepancyYuankun Jiang, Chenglin Li, Wenrui Dai, Junni Zou et al.ICML 2021 · 24 citations
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang et al.ICML 2022 · 78 citations
- Single Episode Policy Transfer in Reinforcement LearningJiachen Yang, Brenden K. Petersen, Hongyuan Zha, Daniel M. FaissolICLR 2020 · 38 citations
- Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single PolicyBram Grooten, Patrick MacAlpine, Kaushik Subramanian, Peter Stone et al.AAAI 2026 · 2 citations
