Maximum State Entropy Exploration using Predecessor and Successor Representations
Arnav Kumar Jain, Lucas Lehnert, Irina Rish, Glen Berseth
摘要
Animals have a developed ability to explore that aids them in important tasks such as locating food, exploring for shelter, and finding misplaced items. These exploration skills necessarily track where they have been so that they can plan for finding items with relative efficiency. Contemporary exploration algorithms often learn a less efficient exploration strategy because they either condition only on the current state or simply rely on making random open-loop exploratory moves. In this work, we propose -Learning, a method to learn efficient exploratory policies by conditioning on past episodic experience to make the next exploratory move. Specifically, -Learning learns an exploration policy that maximizes the entropy of the state visitation distribution of a single trajectory. Furthermore, we demonstrate how variants of the predecessor representation and successor representations can be combined to predict the state visitation entropy. Our experiments demonstrate the efficacy of -Learning to strategically explore the environment and maximize the state coverage with limited samples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- A Distributional Analogue to the Successor RepresentationHarley Wiltzer, Jesse Farebrother, Arthur Gretton, Yunhao Tang 等ICML 2024 · 被引用 11 次
- State Entropy Regularization for Robust Reinforcement LearningYonatan Ashlag, Uri Koren, Mirco Mutti, Esther Derman 等NeurIPS 2025 · 被引用 9 次
- Intention-Conditioned Flow Occupancy ModelsChongyi Zheng, Seohong Park, Sergey Levine, Benjamin EysenbachICLR 2026 · 被引用 9 次
- How to Explore with Belief: State Entropy Maximization in POMDPsRiccardo Zamboni, Duilio Cirino, Marcello Restelli, Mirco MuttiICML 2024 · 被引用 7 次
- Towards Principled Unsupervised Multi-Agent Reinforcement LearningRiccardo Zamboni, Mirco Mutti, Marcello RestelliNeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper16
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel 等ICML 2020 · 被引用 489 次
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair 等ICML 2020 · 被引用 303 次
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 被引用 206 次
- Fast Task Inference with Variational Intrinsic Successor FeaturesSteven Hansen, Will Dabney, André Barreto, David Warde-Farley 等ICLR 2020 · 被引用 176 次
相关 Paper
- Task-Agnostic Exploration via Policy Gradient of a Non-Parametric State Entropy EstimateMirco Mutti, Lorenzo Pratissoli, Marcello RestelliAAAI 2021 · 被引用 62 次
- State Entropy Maximization with Random Encoders for Efficient ExplorationYounggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee 等ICML 2021 · 被引用 158 次
- Exploration by Learning Diverse Skills through Successor State RepresentationsPaul-Antoine Le Tolguenec, Yann Besse, Florent Teichteil-Königsbuch, Dennis Wilson 等NeurIPS 2024
- Toward Efficient Multi-Agent Exploration With Trajectory Entropy MaximizationTianxu Li, Kun ZhuICLR 2025
- A First-Occupancy Representation for Reinforcement LearningTed Moskovitz, Spencer R. Wilson, Maneesh SahaniICLR 2022 · 被引用 18 次
