Stein Variational Goal Generation for adaptive Exploration in Multi-Goal Reinforcement Learning
Nicolas Castanet, Olivier Sigaud, Sylvain Lamprier
Abstract
In multi-goal Reinforcement Learning, an agent can share experience between related training tasks, resulting in better generalization for new tasks at test time. However, when the goal space has discontinuities and the reward is sparse, a majority of goals are difficult to reach. In this context, a curriculum over goals helps agents learn by adapting training tasks to their current capabilities. In this work we propose Stein Variational Goal Generation (SVGG), which samples goals of intermediate difficulty for the agent, by leveraging a learned predictive model of its goal reaching capabilities. The distribution of goals is modeled with particles that are attracted in areas of appropriate difficulty using Stein Variational Gradient Descent. We show that SVGG outperforms state-of-the-art multi-goal Reinforcement Learning methods in terms of success coverage in hard exploration problems, and demonstrate that it is endowed with a useful recovery property when the environment changes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 34fbaa2e-6270-4906-bb3f-85dafc60900cCited by top-tier papers4
- S2AC: Energy-Based Reinforcement Learning with Stein Soft Actor CriticSafa Messaoud, Billel Mokeddem, Zhenghai Xue, Linsey Pang et al.ICLR 2024 · 21 citations
- MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spacesLoris Gaven, Thomas Carta, Clément Romac, Cédric Colas et al.ICML 2025
- Improving Zero-Shot Offline RL via Behavioral Task SamplingNazim Bendib, Nicolas Perrin-Gilbert, Olivier SigaudICML 2026
- Imagine Beyond ! Distributionally Robust Autoencoding for State Space Coverage in Online Reinforcement LearningNicolas Castanet, Olivier Sigaud, Sylvain LamprierNeurIPS 2025
Builds on8
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Reinforcement Learning with Prototypical RepresentationsDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICML 2021 · 262 citations
- Behavior From the Void: Unsupervised Active Pre-TrainingHao Liu, Pieter AbbeelNeurIPS 2021 · 258 citations
- Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement LearningSilviu Pitis, Harris Chan, Stephen Zhao, Bradly C. Stadie et al.ICML 2020 · 145 citations
- Automatic Curriculum Learning through Value DisagreementYunzhi Zhang, Pieter Abbeel, Lerrel PintoNeurIPS 2020 · 132 citations
Related papers
- Stochastic Multiple Target Sampling Gradient DescentHoang Phan, Ngoc Tran, Trung Le, Toan Tran et al.NeurIPS 2022 · 17 citations
- Variational Automatic Curriculum Learning for Sparse-Reward Cooperative Multi-Agent ProblemsJiayu Chen, Yuanxin Zhang, Yuanfan Xu, Huimin Ma et al.NeurIPS 2021 · 48 citations
- Automated curriculum generation through setter-solver interactionsSébastien Racanière, Andrew K. Lampinen, Adam Santoro, David P. Reichert et al.ICLR 2020 · 41 citations
- It Takes Four to Tango: Multiagent Self Play for Automatic Curriculum GenerationYuqing Du, Pieter Abbeel, Aditya GroverICLR 2022 · 20 citations
- Adaptive Procedural Task Generation for Hard-Exploration ProblemsKuan Fang, Yuke Zhu, Silvio Savarese, Li Fei-FeiICLR 2021 · 36 citations
