Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
Dongyoung Kim, Jinwoo Shin, Pieter Abbeel, Younggyo Seo
摘要
A promising technique for exploration is to maximize the entropy of visited state distribution, i.e., state entropy, by encouraging uniform coverage of visited state space. While it has been effective for an unsupervised setup, it tends to struggle in a supervised setup with a task reward, where an agent prefers to visit high-value states to exploit the task reward. Such a preference can cause an imbalance between the distributions of high-value states and low-value states, which biases exploration towards low-value state regions as a result of the state entropy increasing when the distribution becomes more uniform. This issue is exacerbated when high-value states are narrowly distributed within the state space, making it difficult for the agent to complete the tasks. In this paper, we present a novel exploration technique that maximizes the value-conditional state entropy, which separately estimates the state entropies that are conditioned on the value estimates of each state, then maximizes their average. By only considering the visited states with similar value estimates for computing the intrinsic bonus, our method prevents the distribution of low-value states from affecting exploration around high-value states, and vice versa. We demonstrate that the proposed alternative to the state entropy baseline significantly accelerates various reinforcement learning algorithms across a variety of tasks within MiniGrid, DeepMind Control Suite, and Meta-World benchmarks. Source code is available at https://sites.google.com/view/rl-vcse .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- DrM: Mastering Visual Reinforcement Learning through Dormant Ratio MinimizationGuowei Xu, Ruijie Zheng, Yongyuan Liang, Xiyao Wang 等ICLR 2024 · 被引用 53 次
- Deep Bayesian Active Learning for Preference Modeling in Large Language ModelsLuckeciano Carvalho Melo, Panagiotis Tigas, Alessandro Abate, Yarin GalNeurIPS 2024 · 被引用 25 次
- Effective Exploration Based on the Structural Information PrinciplesXianghua Zeng, Hao Peng, Angsheng LiNeurIPS 2024 · 被引用 14 次
- State Entropy Regularization for Robust Reinforcement LearningYonatan Ashlag, Uri Koren, Mirco Mutti, Esther Derman 等NeurIPS 2025 · 被引用 9 次
- How to Explore with Belief: State Entropy Maximization in POMDPsRiccardo Zamboni, Duilio Cirino, Marcello Restelli, Mirco MuttiICML 2024 · 被引用 7 次
它引用的顶会 Paper14
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel 等ICML 2020 · 被引用 489 次
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 被引用 457 次
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo 等ICLR 2020 · 被引用 349 次
- Reinforcement Learning with Prototypical RepresentationsDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICML 2021 · 被引用 262 次
相关 Paper
- State Entropy Maximization with Random Encoders for Efficient ExplorationYounggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee 等ICML 2021 · 被引用 158 次
- Towards Principled Unsupervised Multi-Agent Reinforcement LearningRiccardo Zamboni, Mirco Mutti, Marcello RestelliNeurIPS 2025 · 被引用 5 次
- Unsupervised Reinforcement Learning in Multiple EnvironmentsMirco Mutti, Mattia Mancassola, Marcello RestelliAAAI 2022 · 被引用 30 次
- CEM: Constrained Entropy Maximization for Task-Agnostic Safe ExplorationQisong Yang, Matthijs T. J. SpaanAAAI 2023 · 被引用 24 次
- Rethinking Exploration in Reinforcement Learning with Effective Metric-Based Exploration BonusYiming Wang, Kaiyan Zhao, Furui Liu, Leong Hou UNeurIPS 2024 · 被引用 15 次
