CQM: Curriculum Reinforcement Learning with a Quantized World Model
Seungjae Lee, Daesol Cho, Jonghae Park, H. Jin Kim
Abstract
Recent curriculum Reinforcement Learning (RL) has shown notable progress in solving complex tasks by proposing sequences of surrogate tasks. However, the previous approaches often face challenges when they generate curriculum goals in a high-dimensional space. Thus, they usually rely on manually specified goal spaces. To alleviate this limitation and improve the scalability of the curriculum, we propose a novel curriculum method that automatically defines the semantic goal space which contains vital information for the curriculum process, and suggests curriculum goals over it. To define the semantic goal space, our method discretizes continuous observations via vector quantized-variational autoencoders (VQ-VAE) and restores the temporal relations between the discretized observations by a graph. Concurrently, ours suggests uncertainty and temporal distance-aware curriculum goals that converges to the final goals over the automatically composed goal space. We demonstrate that the proposed method allows efficient explorations in an uninformed environment with raw goal examples only. Also, ours outperforms the state-of-the-art curriculum RL methods on data efficiency and performance, in various goal-reaching tasks even with ego-centric visual inputs. Related Works Curriculum Goal Generation. Although various prior studies [41, 19, 48, 8, 45, 22] have been proposed to solve exploration problems, enabling efficient searching in uninformed environments still
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b76e5fd-d7c9-4a7a-8118-803e1c66938bCited by top-tier papers5
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement LearningLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas KrauseNeurIPS 2025 · 14 citations
- Adversarial Environment Design via Regret-Guided Diffusion ModelsHojun Chung, Junseo Lee, Minsoo Kim, Dohyeong Kim et al.NeurIPS 2024 · 11 citations
- Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement LearningSeungyul Han, Jaebak Hwang, Sanghyeon Lee, Jeongmo KimICLR 2026 · 3 citations
- Periodic Skill DiscoveryJonghae Park, Daesol Cho, Jusuk Lee, Dongseok Shim et al.NeurIPS 2025 · 3 citations
- Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual ForagingBo Wang, Dingwei Tan, Yen-Ling Kuo, Zhaowei Sun et al.CVPR 2025
Builds on21
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm et al.ICLR 2021 · 399 citations
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Reinforcement Learning with Prototypical RepresentationsDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICML 2021 · 262 citations
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 211 citations
Related papers
- Outcome-directed Reinforcement Learning by Uncertainty & Temporal Distance-Aware Curriculum Goal GenerationDaesol Cho, Seungjae Lee, H. Jin KimICLR 2023
- Variational Curriculum Reinforcement Learning for Unsupervised Discovery of SkillsSeongun Kim, Kyowoon Lee, Jaesik ChoiICML 2023 · 17 citations
- Self-Paced Deep Reinforcement LearningPascal Klink, Carlo D'Eramo, Jan Peters, Joni PajarinenNeurIPS 2020 · 83 citations
- Automatic Curriculum Learning through Value DisagreementYunzhi Zhang, Pieter Abbeel, Lerrel PintoNeurIPS 2020 · 132 citations
- Diversify & Conquer: Outcome-directed Curriculum RL via Out-of-Distribution DisagreementDaesol Cho, Seungjae Lee, H. Jin KimNeurIPS 2023 · 4 citations
