CQM: Curriculum Reinforcement Learning with a Quantized World Model
Seungjae Lee, Daesol Cho, Jonghae Park, H. Jin Kim
摘要
Recent curriculum Reinforcement Learning (RL) has shown notable progress in solving complex tasks by proposing sequences of surrogate tasks. However, the previous approaches often face challenges when they generate curriculum goals in a high-dimensional space. Thus, they usually rely on manually specified goal spaces. To alleviate this limitation and improve the scalability of the curriculum, we propose a novel curriculum method that automatically defines the semantic goal space which contains vital information for the curriculum process, and suggests curriculum goals over it. To define the semantic goal space, our method discretizes continuous observations via vector quantized-variational autoencoders (VQ-VAE) and restores the temporal relations between the discretized observations by a graph. Concurrently, ours suggests uncertainty and temporal distance-aware curriculum goals that converges to the final goals over the automatically composed goal space. We demonstrate that the proposed method allows efficient explorations in an uninformed environment with raw goal examples only. Also, ours outperforms the state-of-the-art curriculum RL methods on data efficiency and performance, in various goal-reaching tasks even with ego-centric visual inputs. Related Works Curriculum Goal Generation. Although various prior studies [41, 19, 48, 8, 45, 22] have been proposed to solve exploration problems, enabling efficient searching in uninformed environments still
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement LearningLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas KrauseNeurIPS 2025 · 被引用 14 次
- Adversarial Environment Design via Regret-Guided Diffusion ModelsHojun Chung, Junseo Lee, Minsoo Kim, Dohyeong Kim 等NeurIPS 2024 · 被引用 11 次
- Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement LearningSeungyul Han, Jaebak Hwang, Sanghyeon Lee, Jeongmo KimICLR 2026 · 被引用 3 次
- Periodic Skill DiscoveryJonghae Park, Daesol Cho, Jusuk Lee, Dongseok Shim 等NeurIPS 2025 · 被引用 3 次
- Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual ForagingBo Wang, Dingwei Tan, Yen-Ling Kuo, Zhaowei Sun 等CVPR 2025
它引用的顶会 Paper21
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm 等ICLR 2021 · 被引用 399 次
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair 等ICML 2020 · 被引用 303 次
- Reinforcement Learning with Prototypical RepresentationsDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICML 2021 · 被引用 262 次
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 被引用 211 次
相关 Paper
- Outcome-directed Reinforcement Learning by Uncertainty & Temporal Distance-Aware Curriculum Goal GenerationDaesol Cho, Seungjae Lee, H. Jin KimICLR 2023
- Variational Curriculum Reinforcement Learning for Unsupervised Discovery of SkillsSeongun Kim, Kyowoon Lee, Jaesik ChoiICML 2023 · 被引用 17 次
- Self-Paced Deep Reinforcement LearningPascal Klink, Carlo D'Eramo, Jan Peters, Joni PajarinenNeurIPS 2020 · 被引用 83 次
- Automatic Curriculum Learning through Value DisagreementYunzhi Zhang, Pieter Abbeel, Lerrel PintoNeurIPS 2020 · 被引用 132 次
- Diversify & Conquer: Outcome-directed Curriculum RL via Out-of-Distribution DisagreementDaesol Cho, Seungjae Lee, H. Jin KimNeurIPS 2023 · 被引用 4 次
