NetHack is Hard to Hack
Ulyana Piterbarg, Lerrel Pinto, Rob Fergus
摘要
Neural policy learning methods have achieved remarkable results in various control problems, ranging from Atari games to simulated locomotion. However, these methods struggle in long-horizon tasks, especially in open-ended environments with multi-modal observations, such as the popular dungeon-crawler game, NetHack. Intriguingly, the NeurIPS 2021 NetHack Challenge revealed that symbolic agents outperformed neural approaches by over four times in median game score. In this paper, we delve into the reasons behind this performance gap and present an extensive study on neural policy learning for NetHack. To conduct this study, we analyze the winning symbolic agent, extending its codebase to track internal strategy selection in order to generate one of the largest available demonstration datasets. Utilizing this dataset, we examine (i) the advantages of an action hierarchy; (ii) enhancements in neural architecture; and (iii) the integration of reinforcement learning with imitation learning. Our investigations produce a state-of-the-art neural agent that surpasses previous fully neural policies by 127% in offline settings and 25% in online settings on median game score. However, we also demonstrate that mere scaling is insufficient to bridge the performance gap with the best symbolic models or even the top human players.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Motif: Intrinsic Motivation from Artificial Intelligence FeedbackMartin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu 等ICLR 2024 · 被引用 97 次
- Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation ProblemMaciej Wolczyk, Bartlomiej Cupial, Mateusz Ostaszewski, Michal Bortkiewicz 等ICML 2024 · 被引用 29 次
- diff History for Neural Language AgentsUlyana Piterbarg, Lerrel Pinto, Rob FergusICML 2024 · 被引用 4 次
- RefactorBench: Evaluating Stateful Reasoning in Language Agents Through CodeDhruv Gautam, Spandan Garg, Jinu Jang, Neel Sundaresan 等ICLR 2025
- MaestroMotif: Skill Design from Artificial Intelligence FeedbackMartin Klissarov, Mikael Henaff, Roberta Raileanu, Shagun Sodhani 等ICLR 2025
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals 等ICML 2021 · 被引用 1,399 次
相关 Paper
- Learning Symbolic Rules over Abstract Meaning Representations for Textual Reinforcement LearningSubhajit Chaudhury, Sarathkrishna Swaminathan, Daiki Kimura, Prithviraj Sen 等ACL 2023 · 被引用 4 次
- Scalable Option Learning in High-Throughput EnvironmentsMikael Henaff, Scott Fujimoto, Michael Matthews, Michael RabbatICML 2026 · 被引用 5 次
- Discovering symbolic policies with deep reinforcement learningMikel Landajuela, Brenden K. Petersen, Sookyung Kim, Cláudio P. Santiago 等ICML 2021 · 被引用 118 次
- BlendRL: A Framework for Merging Symbolic and Neural Policy LearningHikaru Shindo, Quentin Delfosse, Devendra Singh Dhami, Kristian KerstingICLR 2025
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu 等NeurIPS 2020 · 被引用 251 次
