Discovering Hierarchical Achievements in Reinforcement Learning via Contrastive Learning
Seungyong Moon, Junyoung Yeom, Bumsoo Park, Hyun Oh Song
Abstract
Discovering achievements with a hierarchical structure in procedurally generated environments presents a significant challenge. This requires an agent to possess a broad range of abilities, including generalization and long-term reasoning. Many prior methods have been built upon model-based or hierarchical approaches, with the belief that an explicit module for long-term planning would be advantageous for learning hierarchical dependencies. However, these methods demand an excessive number of environment interactions or large model sizes, limiting their practicality. In this work, we demonstrate that proximal policy optimization (PPO), a simple yet versatile model-free algorithm, outperforms previous methods when optimized with recent implementation practices. Moreover, we find that the PPO agent can predict the next achievement to be unlocked to some extent, albeit with limited confidence. Based on this observation, we introduce a novel contrastive learning method, called achievement distillation, which strengthens the agent's ability to predict the next achievement. Our method exhibits a strong capacity for discovering hierarchical achievements and shows state-of-the-art performance on the challenging Crafter environment in a sample-efficient manner while utilizing fewer model parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ed3e057-8e84-4eab-aa86-c610db4eb2d7Cited by top-tier papers4
- Pre-Trained Multi-Goal Transformers with Prompt Optimization for Efficient Online AdaptationHaoqi Yuan, Yuhui Fu, Feiyang Xie, Zongqing LuNeurIPS 2024 · 5 citations
- Identifiable Token Correspondence for World ModelsYoungin Kim, Ray Sun, Inho Kim, Bumsoo Park et al.ICML 2026
- Counterfactual Planning for Generalizable Agents' ActionsJiarun Fu, Lizhong Ding, Qiuning Wei, Yuhan Guo et al.AAAI 2026
- Improving Transformer World Models for Data-Efficient RLAntoine Dedieu, Joseph Ortiz, Xinghua Lou, Carter Wendelken et al.ICML 2025
Builds on26
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesDenis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos et al.AAAI 2021 · 506 citations
Related papers
- Learning Achievement Structure for Structured Exploration in Domains with Sparse RewardZihan Zhou, Animesh GargICLR 2023
- Benchmarking the Spectrum of Agent CapabilitiesDanijar HafnerICLR 2022 · 193 citations
- Polychromic Objectives for Reinforcement LearningJubayer Ibn Hamid, Ifdita Hasan Orney, Ellen Xu, Chelsea Finn et al.ICLR 2026 · 9 citations
- REBEL: Reinforcement Learning via Regressing Relative RewardsZhaolin Gao, Jonathan D. Chang, Wenhao Zhan, Owen Oertell et al.NeurIPS 2024 · 82 citations
- Guided Exploration with Proximal Policy Optimization using a Single DemonstrationGabriele Libardi, Gianni De Fabritiis, Sebastian DittertICML 2021 · 32 citations
