DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
Leander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas Krause
Abstract
Sparse-reward reinforcement learning (RL) can model a wide range of highly complex tasks. Solving sparse-reward tasks is RL's core premise, requiring efficient exploration coupled with long-horizon credit assignment, and overcoming these challenges is key for building self-improving agents with superhuman ability. Prior work commonly explores with the objective of solving many sparse-reward tasks, making exploration of individual high-dimensional, long-horizon tasks intractable. We argue that solving such challenging tasks requires solving simpler tasks that are relevant to the target task, i.e., whose achieval will teach the agent skills required for solving the target task. We demonstrate that this sense of direction, necessary for effective exploration, can be extracted from existing RL algorithms, without leveraging any prior information. To this end, we propose a method for directed sparse-reward goal-conditioned very long-horizon RL (DISCOVER), which selects exploratory goals in the direction of the target task. We connect DISCOVER to principled exploration in bandits, formally bounding the time until the target task becomes achievable in terms of the agent's initial distance to the target, but independent of the volume of the space of all tasks. We then perform a thorough evaluation in high-dimensional environments. We find that the directed goal selection of DISCOVER solves exploration problems that are beyond the reach of prior state-of-the-art exploration methods in RL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0f59ca6-0373-4383-a9a0-98def156ea2cCited by top-tier papers3
- Test-time Offline Reinforcement Learning on Goal-related ExperienceMarco Bagatella, Mert Albaba, Jonas Hübotter, Georg Martius et al.ICML 2026 · 7 citations
- Specialization after Generalization: Towards Understanding Test-Time Training in Foundation ModelsJonas Hübotter, Patrik Wolf, Aleksandr Shevchenko, Dennis Jüni et al.ICLR 2026 · 6 citations
- Reinforcement Learning via Self-DistillationJonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann et al.ICML 2026
Builds on36
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller et al.ICML 2020 · 1,220 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen et al.NeurIPS 2020 · 362 citations
- Absolute Zero: Reinforced Self-play Reasoning with Zero DataAndrew Zhao, Yiran Wu, Tong Wu, Quentin Xu et al.NeurIPS 2025 · 361 citations
Related papers
- Subspace-Aware Exploration for Sparse-Reward Multi-Agent TasksPei Xu, Junge Zhang, Qiyue Yin, Chao Yu et al.AAAI 2023 · 15 citations
- Active Hierarchical Exploration with Stable Subgoal Representation LearningSiyuan Li, Jin Zhang, Jianhao Wang, Yang Yu et al.ICLR 2022 · 28 citations
- Hierarchical Reinforcement Learning with Targeted Causal InterventionsMohammadsadegh Khorasani, Saber Salehkaleybar, Negar Kiyavash, Matthias GrossglauserICML 2025
- Learning Guidance Rewards with Trajectory-space SmoothingTanmay Gangwani, Yuan Zhou, Jian PengNeurIPS 2020 · 46 citations
- Deep Hierarchical Planning from PixelsDanijar Hafner, Kuang-Huei Lee, Ian Fischer, Pieter AbbeelNeurIPS 2022 · 153 citations
