Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning
Silviu Pitis, Harris Chan, Stephen Zhao, Bradly C. Stadie, Jimmy Ba
Abstract
What goals should a multi-goal reinforcement learning agent pursue during training in longhorizon tasks? When the desired (test time) goal distribution is too distant to offer a useful learning signal, we argue that the agent should not pursue unobtainable goals. Instead, it should set its own intrinsic goals that maximize the entropy of the historical achieved goal distribution. We propose to optimize this objective by having the agent pursue past achieved goals in sparsely explored areas of the goal space, which focuses exploration on the frontier of the achievable goal set. We show that our strategy achieves an order of magnitude better sample efficiency than the prior state of the art on long-horizon multi-goal tasks including maze navigation and block stacking. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 417d99c2-1431-4df9-bce2-678cdd9e58f2Cited by top-tier papers58
- Behavior From the Void: Unsupervised Active Pre-TrainingHao Liu, Pieter AbbeelNeurIPS 2021 · 258 citations
- Goal-Conditioned Reinforcement Learning with Imagined SubgoalsElliot Chane-Sane, Cordelia Schmid, Ivan LaptevICML 2021 · 183 citations
- Planning Goals for ExplorationEdward S. Hu, Richard Chang, Oleh Rybkin, Dinesh JayaramanICLR 2023 · 152 citations
- Counterfactual Data Augmentation using Locally Factored DynamicsSilviu Pitis, Elliot Creager, Animesh GargNeurIPS 2020 · 126 citations
- BYOL-Explore: Exploration by Bootstrapped PredictionZhaohan Guo, Shantanu Thakoor, Miruna Pislar, Bernardo Ávila Pires et al.NeurIPS 2022 · 104 citations
Builds on5
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Fast Task Inference with Variational Intrinsic Successor FeaturesSteven Hansen, Will Dabney, André Barreto, David Warde-Farley et al.ICLR 2020 · 176 citations
- Automatic Curriculum Learning through Value DisagreementYunzhi Zhang, Pieter Abbeel, Lerrel PintoNeurIPS 2020 · 132 citations
- Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill DiscoveryKristian Hartikainen, Xinyang Geng, Tuomas Haarnoja, Sergey LevineICLR 2020 · 94 citations
- Weakly-Supervised Reinforcement Learning for Controllable BehaviorLisa Lee, Ben Eysenbach, Ruslan Salakhutdinov, Shixiang Shane Gu et al.NeurIPS 2020 · 28 citations
Related papers
- Stein Variational Goal Generation for adaptive Exploration in Multi-Goal Reinforcement LearningNicolas Castanet, Olivier Sigaud, Sylvain LamprierICML 2023 · 6 citations
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement LearningLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas KrauseNeurIPS 2025 · 14 citations
- Enhancing Exploration and Exploitation in Hierarchical Reinforcement Learning with Subgoal Graph LearningYibo Zhang, Dengpeng XingAAAI 2026
- Hierarchical Reinforcement Learning with Targeted Causal InterventionsMohammadsadegh Khorasani, Saber Salehkaleybar, Negar Kiyavash, Matthias GrossglauserICML 2025
- Go Beyond Imagination: Maximizing Episodic Reachability with World ModelsYao Fu, Run Peng, Honglak LeeICML 2023 · 1 citation
