Ask Your Humans: Using Human Instructions to Improve Generalization in Reinforcement Learning
Valerie Chen, Abhinav Gupta, Kenneth Marino
摘要
Complex, multi-task problems have proven to be difficult to solve efficiently in a sparse-reward reinforcement learning setting. In order to be sample efficient, multi-task learning requires reuse and sharing of low-level policies. To facilitate the automatic decomposition of hierarchical tasks, we propose the use of step-by-step human demonstrations in the form of natural language instructions and action trajectories. We introduce a dataset of such demonstrations in a crafting-based grid world. Our model consists of a high-level language generator and low-level policy, conditioned on language. We find that human demonstrations help solve the most complex tasks. We also find that incorporating natural language allows the model to generalize to unseen tasks in a zero-shot setting and to learn quickly from a few demonstrations. Generalization is not only reflected in the actions of the agent, but also in the generated natural language instructions in unseen tasks. Our approach also gives our trained agent interpretable behaviors because it is able to generate a sequence of high-level descriptions of its actions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- LISA: Learning Interpretable Skill Abstractions from LanguageDivyansh Garg, Skanda Vaidyanath, Kuno Kim, Jiaming Song 等NeurIPS 2022 · 被引用 43 次
- Distilling Internet-Scale Vision-Language Models into Embodied AgentsTheodore R. Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus 等ICML 2023 · 被引用 36 次
- How to talk so AI will learn: Instructions, descriptions, and autonomyTheodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Tom Griffiths 等NeurIPS 2022 · 被引用 30 次
- Perceiving the World: Question-guided Reinforcement Learning for Text-based GamesYunqiu Xu, Meng Fang, Ling Chen, Yali Du 等ACL 2022 · 被引用 22 次
- Improving Long-Horizon Imitation through Instruction PredictionJoey Hejna, Pieter Abbeel, Lerrel PintoAAAI 2023 · 被引用 12 次
它引用的顶会 Paper1
相关 Paper
- Skill Induction and Planning with Latent LanguagePratyusha Sharma, Antonio Torralba, Jacob AndreasACL 2022 · 被引用 127 次
- Learning Compositional Tasks from Language InstructionsLajanugen Logeswaran, Wilka Carvalho, Honglak LeeAAAI 2023 · 被引用 4 次
- Learning with Language-Guided State AbstractionsAndi Peng, Ilia Sucholutsky, Belinda Z. Li, Theodore R. Sumers 等ICLR 2024 · 被引用 20 次
- Zero-Shot Reward Specification via Grounded Natural LanguageParsa Mahmoudieh, Deepak Pathak, Trevor DarrellICML 2022 · 被引用 69 次
- RLZero: Direct Policy Inference from Language Without In-Domain SupervisionHarshit Sikchi, Siddhant Agarwal, Pranaya Jajoo, Samyak Parajuli 等NeurIPS 2025 · 被引用 8 次
