Improving Policy Learning via Language Dynamics Distillation
Victor Zhong, Jesse Mu, Luke Zettlemoyer, Edward Grefenstette, Tim Rocktäschel
摘要
Recent work has shown that augmenting environments with language descriptions improves policy learning. However, for environments with complex language abstractions, learning how to ground language to observations is difficult due to sparse, delayed rewards. We propose Language Dynamics Distillation (LDD), which pretrains a model to predict environment dynamics given demonstrations with language descriptions, and then fine-tunes these language-aware pretrained representations via reinforcement learning (RL). In this way, the model is trained to both maximize expected reward and retain knowledge about how language relates to environment dynamics. On SILG, a benchmark of five tasks with language descriptions that evaluate distinct generalization challenges on unseen environments (NetHack, ALFWorld, RTFM, Messenger, and Touchdown), LDD outperforms tabula-rasa RL, VAE pretraining, and methods that learn from unlabeled demonstrations in inverse RL and reward shaping with pretrained experts. In our analyses, we show that language descriptions in demonstrations improve sample-efficiency and generalization across environments, and that dynamics modelling with expert demonstrations is more effective than with non-experts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Learning to Model the World With LanguageJessy Lin, Yuqing Du, Olivia Watkins, Danijar Hafner 等ICML 2024 · 被引用 76 次
- Distilling Internet-Scale Vision-Language Models into Embodied AgentsTheodore R. Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus 等ICML 2023 · 被引用 36 次
- Language Control Diffusion: Efficiently Scaling through Space, Time, and TasksEdwin Zhang, Yujie Lu, Shinda Huang, William Yang Wang 等ICLR 2024 · 被引用 34 次
- Semantic HELM: A Human-Readable Memory for Reinforcement LearningFabian Paischer, Thomas Adler, Markus Hofmarcher, Sepp HochreiterNeurIPS 2023 · 被引用 21 次
- diff History for Neural Language AgentsUlyana Piterbarg, Lerrel Pinto, Rob FergusICML 2024 · 被引用 4 次
它引用的顶会 Paper8
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk 等ICLR 2021 · 被引用 819 次
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu 等NeurIPS 2020 · 被引用 251 次
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie 等EMNLP 2020 · 被引用 208 次
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 被引用 198 次
- Semantic Exploration from Language Abstractions and Pretrained RepresentationsAllison C. Tam, Neil C. Rabinowitz, Andrew K. Lampinen, Nicholas A. Roy 等NeurIPS 2022 · 被引用 85 次
相关 Paper
- SILG: The Multi-domain Symbolic Interactive Language Grounding BenchmarkVictor Zhong, Austin W. Hanjie, Sida I. Wang, Karthik Narasimhan 等NeurIPS 2021 · 被引用 23 次
- Bridging Environments and Language with Rendering Functions and Vision-Language ModelsThéo Cachet, Christopher R. Dance, Olivier SigaudICML 2024 · 被引用 1 次
- Disentangled Robot Learning via Separate Forward and Inverse Dynamics PretrainingWenyao Zhang, Bozhou Zhang, Zekun Qi, Wenjun Zeng 等ICLR 2026 · 被引用 18 次
- Grounding Language to Entities and Dynamics for Generalization in Reinforcement LearningAustin W. Hanjie, Victor Zhong, Karthik NarasimhanICML 2021 · 被引用 60 次
- Conceptual Reinforcement Learning for Language-Conditioned TasksShaohui Peng, Xing Hu, Rui Zhang, Jiaming Guo 等AAAI 2023 · 被引用 12 次
