Skill Induction and Planning with Latent Language
Pratyusha Sharma, Antonio Torralba, Jacob Andreas
Abstract
We present a framework for learning hierarchical policies from demonstrations, using sparse natural language annotations to guide the discovery of reusable skills for autonomous decision-making. We formulate a generative model of action sequences in which goals generate sequences of high-level subtask descriptions, and these descriptions generate sequences of low-level actions. We describe how to train this model using primarily unannotated demonstrations by parsing demonstrations into sequences of named high-level subtasks, using only a small number of seed annotations to ground language in action. In trained models, natural language commands index a combinatorial library of skills; agents can use these skills to plan by generating high-level instruction sequences tailored to novel goals. We evaluate this approach in the ALFRED household simulation environment, providing natural language annotations for only 10% of demonstrations. It achieves task completion rates comparable to state-of-the-art models (outperforming several recent methods with access to ground-truth plans during training and evaluation) while providing structured and human-readable high-level plans. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c38143d6-edab-48e5-8ab9-474c9b3dff6fCited by top-tier papers40
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
- Large Language Models as Commonsense Knowledge for Large-Scale Task PlanningZirui Zhao, Wee Sun Lee, David HsuNeurIPS 2023 · 423 citations
- Pre-Trained Language Models for Interactive Decision-MakingShuang Li, Xavier Puig, Chris Paxton, Yilun Du et al.NeurIPS 2022 · 341 citations
Builds on6
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
- Episodic Transformer for Vision-and-Language NavigationAlexander Pashevich, Cordelia Schmid, Chen SunICCV 2021 · 228 citations
- FILM: Following Instructions in Language with Modular MethodsSo Yeon Min, Devendra Singh Chaplot, Pradeep Kumar Ravikumar, Yonatan Bisk et al.ICLR 2022 · 189 citations
- Leveraging Language to Learn Program Abstractions and Search HeuristicsCatherine Wong, Kevin Ellis, Joshua B. Tenenbaum, Jacob AndreasICML 2021 · 59 citations
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday TasksMohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk et al.CVPR 2020
Related papers
- Language-guided Skill Learning with Temporal Variational InferenceHaotian Fu, Pratyusha Sharma, Elias Stengel-Eskin, George Konidaris et al.ICML 2024 · 11 citations
- Ask Your Humans: Using Human Instructions to Improve Generalization in Reinforcement LearningValerie Chen, Abhinav Gupta, Kenneth MarinoICLR 2021 · 6 citations
- Learning Grounded Action Abstractions from LanguageLionel Wong, Jiayuan Mao, Pratyusha Sharma, Zachary S. Siegel et al.ICLR 2024 · 7 citations
- LISA: Learning Interpretable Skill Abstractions from LanguageDivyansh Garg, Skanda Vaidyanath, Kuno Kim, Jiaming Song et al.NeurIPS 2022 · 43 citations
- Learning Planning Abstractions from LanguageWeiyu Liu, Geng Chen, Joy Hsu, Jiayuan Mao et al.ICLR 2024 · 6 citations
