Task-Agnostic Exploration via Policy Gradient of a Non-Parametric State Entropy Estimate
Mirco Mutti, Lorenzo Pratissoli, Marcello Restelli
摘要
In a reward-free environment, what is a suitable intrinsic objective for an agent to pursue so that it can learn an optimal task-agnostic exploration policy? In this paper, we argue that the entropy of the state distribution induced by finite-horizon trajectories is a sensible target. Especially, we present a novel and practical policy-search algorithm, Maximum Entropy POLicy optimization (MEPOL), to learn a policy that maximizes a non-parametric, -nearest neighbors estimate of the state distribution entropy. In contrast to known methods, MEPOL is completely model-free as it requires neither to estimate the state distribution of any policy nor to model transition dynamics. Then, we empirically show that MEPOL allows learning a maximum-entropy exploration policy in high-dimensional, continuous-control domains, and how this policy facilitates learning meaningful reward-based tasks downstream.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Reinforcement Learning with Prototypical RepresentationsDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICML 2021 · 被引用 262 次
- METRA: Scalable Unsupervised RL with Metric-Aware AbstractionSeohong Park, Oleh Rybkin, Sergey LevineICLR 2024 · 被引用 83 次
- The Importance of Non-Markovianity in Maximum State Entropy ExplorationMirco Mutti, Riccardo De Santi, Marcello RestelliICML 2022 · 被引用 45 次
- Fast Rates for Maximum Entropy ExplorationDaniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines 等ICML 2023 · 被引用 34 次
- Challenging Common Assumptions in Convex Reinforcement LearningMirco Mutti, Riccardo De Santi, Piersilvio De Bartolomeis, Marcello RestelliNeurIPS 2022 · 被引用 31 次
它引用的顶会 Paper5
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair 等ICML 2020 · 被引用 303 次
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 被引用 226 次
- Kinematic State Abstraction and Provably Efficient Rich-Observation Reinforcement LearningDipendra Misra, Mikael Henaff, Akshay Krishnamurthy, John LangfordICML 2020 · 被引用 158 次
- Task-agnostic Exploration in Reinforcement LearningXuezhou Zhang, Yuzhe Ma, Adish SinglaNeurIPS 2020 · 被引用 56 次
- An Intrinsically-Motivated Approach for Learning Highly Exploring and Fast Mixing PoliciesMirco Mutti, Marcello RestelliAAAI 2020 · 被引用 31 次
相关 Paper
- CEM: Constrained Entropy Maximization for Task-Agnostic Safe ExplorationQisong Yang, Matthijs T. J. SpaanAAAI 2023 · 被引用 24 次
- Towards Principled Unsupervised Multi-Agent Reinforcement LearningRiccardo Zamboni, Mirco Mutti, Marcello RestelliNeurIPS 2025 · 被引用 5 次
- Exploration by Maximizing Renyi Entropy for Reward-Free RL FrameworkChuheng Zhang, Yuanying Cai, Longbo Huang, Jian LiAAAI 2021 · 被引用 48 次
- Toward Efficient Multi-Agent Exploration With Trajectory Entropy MaximizationTianxu Li, Kun ZhuICLR 2025
- Scalable Online Exploration via CoverabilityPhilip Amortila, Dylan J. Foster, Akshay KrishnamurthyICML 2024 · 被引用 10 次
