Learning About Progress From Experts
Jake Bruce, Ankit Anand, Bogdan Mazoure, Rob Fergus
Abstract
Many important tasks involve some notion of long-term progress in multiple phases: e.g. to clean a shelf it must be cleared of items, cleaning products applied, and then the items placed back on the shelf. In this work, we explore the use of expert demonstrations in long-horizon tasks to learn a monotonically increasing function that summarizes progress. This function can then be used to aid agent exploration in environments with sparse rewards. As a case study we consider the NetHack environment, which requires long-term progress at a variety of scales and is far from being solved by existing approaches. In this environment, we demonstrate that by learning a model of long-term progress from expert data containing only observations, we can achieve efficient exploration in challenging sparse tasks, well beyond what is possible with current state-of-the-art approaches. We have made the curated gameplay dataset used in this work available at https://github.com/deepmind/nao_top10 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66dbc94d-c6f5-4d05-9333-9044092206c2Cited by top-tier papers10
- Motif: Intrinsic Motivation from Artificial Intelligence FeedbackMartin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu et al.ICLR 2024 · 97 citations
- GROOT: Learning to Follow Instructions by Watching Gameplay VideosShaofei Cai, Bowei Zhang, Zihao Wang, Xiaojian Ma et al.ICLR 2024 · 43 citations
- Code as Reward: Empowering Reinforcement Learning with VLMsDavid Venuto, Mohammad Sami Nur Islam, Martin Klissarov, Doina Precup et al.ICML 2024 · 29 citations
- Pre-Training Goal-based Models for Sample-Efficient Reinforcement LearningHaoqi Yuan, Zhancun Mu, Feiyang Xie, Zongqing LuICLR 2024 · 26 citations
- Learning from Visual Observation via Offline Pretrained State-to-Go TransformerBohan Zhou, Ke Li, Jiechuan Jiang, Zongqing LuNeurIPS 2023 · 17 citations
Builds on8
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu et al.NeurIPS 2020 · 251 citations
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 198 citations
- State Entropy Maximization with Random Encoders for Efficient ExplorationYounggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee et al.ICML 2021 · 158 citations
- Reinforcement Learning with Action-Free Pre-Training from VideosYounggyo Seo, Kimin Lee, Stephen James, Pieter AbbeelICML 2022 · 150 citations
Related papers
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement LearningLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas KrauseNeurIPS 2025 · 14 citations
- Skill-based Meta-Reinforcement LearningTaewook Nam, Shao-Hua Sun, Karl Pertsch, Sung Ju Hwang et al.ICLR 2022 · 55 citations
- Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model LearningAdrià López Escoriza, Nicklas Hansen, Stone Tao, Tongzhou Mu et al.ICML 2025
- DeepSynth: Automata Synthesis for Automatic Task Segmentation in Deep Reinforcement LearningMohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham et al.AAAI 2021 · 62 citations
- TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal DistanceYuyang Liu, Chuan Wen, Yihang Hu, Dinesh Jayaraman et al.ICML 2026 · 7 citations
