Unsupervised Reinforcement Learning in Multiple Environments
Mirco Mutti, Mattia Mancassola, Marcello Restelli
摘要
Several recent works have been dedicated to unsupervised reinforcement learning in a single environment, in which a policy is first pre-trained with unsupervised interactions, and then fine-tuned towards the optimal policy for several downstream supervised tasks defined over the same environment. Along this line, we address the problem of unsupervised reinforcement learning in a class of multiple environments, in which the policy is pre-trained with interactions from the whole class, and then fine-tuned for several tasks in any environment of the class. Notably, the problem is inherently multi-objective as we can trade off the pre-training objective between environments in many ways. In this work, we foster an exploration strategy that is sensitive to the most adverse cases within the class. Hence, we cast the exploration problem as the maximization of the mean of a critical percentile of the state visitation entropy induced by the exploration strategy over the class of environments. Then, we present a policy gradient algorithm, alphaMEPOL, to optimize the introduced objective through mediated interactions with the class. Finally, we empirically demonstrate the ability of the algorithm in learning to explore challenging classes of continuous environments and we show that reinforcement learning greatly benefits from the pre-trained exploration strategy w.r.t. learning from scratch.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- The Importance of Non-Markovianity in Maximum State Entropy ExplorationMirco Mutti, Riccardo De Santi, Marcello RestelliICML 2022 · 被引用 45 次
- Fast Rates for Maximum Entropy ExplorationDaniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines 等ICML 2023 · 被引用 34 次
- Accelerating Reinforcement Learning with Value-Conditional State Entropy ExplorationDongyoung Kim, Jinwoo Shin, Pieter Abbeel, Younggyo SeoNeurIPS 2023 · 被引用 34 次
- Challenging Common Assumptions in Convex Reinforcement LearningMirco Mutti, Riccardo De Santi, Piersilvio De Bartolomeis, Marcello RestelliNeurIPS 2022 · 被引用 31 次
- Maximum State Entropy Exploration using Predecessor and Successor RepresentationsArnav Kumar Jain, Lucas Lehnert, Irina Rish, Glen BersethNeurIPS 2023 · 被引用 27 次
它引用的顶会 Paper11
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze 等ICLR 2020 · 被引用 315 次
- Reinforcement Learning with Prototypical RepresentationsDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICML 2021 · 被引用 262 次
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 被引用 226 次
- State Entropy Maximization with Random Encoders for Efficient ExplorationYounggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee 等ICML 2021 · 被引用 158 次
- RL for Latent MDPs: Regret Guarantees and a Lower BoundJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorNeurIPS 2021 · 被引用 91 次
相关 Paper
- Towards Principled Unsupervised Multi-Agent Reinforcement LearningRiccardo Zamboni, Mirco Mutti, Marcello RestelliNeurIPS 2025 · 被引用 5 次
- Task-Agnostic Exploration via Policy Gradient of a Non-Parametric State Entropy EstimateMirco Mutti, Lorenzo Pratissoli, Marcello RestelliAAAI 2021 · 被引用 62 次
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang 等ICML 2022 · 被引用 78 次
- Unsupervised Learning of Efficient Exploration: Pre-training Adaptive Policies via Self-Imposed GoalsOctavio PappalardoICLR 2026
- SEMDICE: Off-policy State Entropy Maximization via Stationary Distribution Correction EstimationJongmin Lee, Meiqi Sun, Pieter AbbeelICLR 2025
