Outcome-directed Reinforcement Learning by Uncertainty & Temporal Distance-Aware Curriculum Goal Generation
Daesol Cho, Seungjae Lee, H. Jin Kim
Abstract
Current reinforcement learning (RL) often suffers when solving a challenging exploration problem where the desired outcomes or high rewards are rarely observed. Even though curriculum RL, a framework that solves complex tasks by proposing a sequence of surrogate tasks, shows reasonable results, most of the previous works still have difficulty in proposing curriculum due to the absence of a mechanism for obtaining calibrated guidance to the desired outcome state without any prior domain knowledge. To alleviate it, we propose an uncertainty & temporal distance-aware curriculum goal generation method for the outcome-directed RL via solving a bipartite matching problem. It could not only provide precisely calibrated guidance of the curriculum to the desired outcome states but also bring much better sample efficiency and geometry-agnostic curriculum goal proposal capability compared to previous curriculum RL methods. We demonstrate that our algorithm significantly outperforms these prior methods in a variety of challenging navigation tasks and robotic manipulation tasks in a quantitative and qualitative way.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5557122-d31c-4b2d-8037-9bcd68322476Cited by top-tier papers10
- Domain Randomization via Entropy MaximizationGabriele Tiboni, Pascal Klink, Jan Peters, Tatiana Tommasi et al.ICLR 2024 · 24 citations
- Can Pre-Trained Text-to-Image Models Generate Visual Goals for Reinforcement Learning?Jialu Gao, Kaizhe Hu, Guowei Xu, Huazhe XuNeurIPS 2023 · 23 citations
- CQM: Curriculum Reinforcement Learning with a Quantized World ModelSeungjae Lee, Daesol Cho, Jonghae Park, H. Jin KimNeurIPS 2023 · 18 citations
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement LearningLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas KrauseNeurIPS 2025 · 14 citations
- Diffusion-based Curriculum Reinforcement LearningErdi Sayar, Giovanni Iacca, Ozgur S. Oguz, Alois KnollNeurIPS 2024 · 12 citations
Builds on17
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Reinforcement Learning with Prototypical RepresentationsDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICML 2021 · 262 citations
- Behavior From the Void: Unsupervised Active Pre-TrainingHao Liu, Pieter AbbeelNeurIPS 2021 · 258 citations
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 211 citations
Related papers
- Automatic Curriculum Learning through Value DisagreementYunzhi Zhang, Pieter Abbeel, Lerrel PintoNeurIPS 2020 · 132 citations
- Diversify & Conquer: Outcome-directed Curriculum RL via Out-of-Distribution DisagreementDaesol Cho, Seungjae Lee, H. Jin KimNeurIPS 2023 · 4 citations
- Curriculum Reinforcement Learning via Constrained Optimal TransportPascal Klink, Haoyi Yang, Carlo D'Eramo, Jan Peters et al.ICML 2022 · 44 citations
- Variational Curriculum Reinforcement Learning for Unsupervised Discovery of SkillsSeongun Kim, Kyowoon Lee, Jaesik ChoiICML 2023 · 17 citations
- Self-Paced Deep Reinforcement LearningPascal Klink, Carlo D'Eramo, Jan Peters, Joni PajarinenNeurIPS 2020 · 83 citations
