Using Both Demonstrations and Language Instructions to Efficiently Learn Robotic Tasks
Albert Yu, Raymond J. Mooney
Abstract
Demonstrations and natural language instructions are two common ways to specify and teach robots novel tasks. However, for many complex tasks, a demonstration or language instruction alone contains ambiguities, preventing tasks from being specified clearly. In such cases, a combination of both a demonstration and an instruction more concisely and effectively conveys the task to the robot than either modality alone. To instantiate this problem setting, we train a single multi-task policy on a few hundred challenging robotic pick-and-place tasks and propose DeL-TaCo (Joint Demo-Language Task Conditioning), a method for conditioning a robotic policy on task embeddings comprised of two components: a visual demonstration and a language instruction. By allowing these two modalities to mutually disambiguate and clarify each other during novel task specification, DeL-TaCo (1) substantially decreases the teacher effort needed to specify a new task and (2) achieves better generalization performance on novel objects and instructions over previous task-conditioning methods. To our knowledge, this is the first work to show that simultaneously conditioning a multi-task robotic manipulation policy on both demonstration and language embeddings improves sample efficiency and generalization over conditioning on either modality alone. See additional materials at https://deltaco-robot.github.io .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Diffusion Model is an Effective Planner and Data Synthesizer for Multi-Task Reinforcement LearningHaoran He, Chenjia Bai, Kang Xu, Zhuoran Yang et al.NeurIPS 2023 · 165 citations
- Large Language Models as Generalizable Policies for Embodied TasksAndrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure et al.ICLR 2024 · 114 citations
- FGPrompt: Fine-grained Goal Prompting for Image-goal NavigationXinyu Sun, Peihao Chen, Jugang Fan, Jian Chen et al.NeurIPS 2023 · 41 citations
- Scaling Context-Aware Task Assistants that Learn from Demonstration and Adapt through Mixed-Initiative DialogueRiku Arakawa, Prasoon Patidar, Will Page, Jill Lehman et al.UIST 2025 · 3 citations
- How to Solve Contextual Goal-Oriented Problems with Offline Datasets?Ying Fan, Jingling Li, Adith Swaminathan, Aditya Modi et al.NeurIPS 2024 · 1 citation
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Multi-Task Reinforcement Learning with Soft ModularizationRuihan Yang, Huazhe Xu, Yi Wu, Xiaolong WangNeurIPS 2020 · 247 citations
- Multi-Task Reinforcement Learning with Context-based RepresentationsShagun Sodhani, Amy Zhang, Joelle PineauICML 2021 · 241 citations
Related papers
- Language-Conditioned Imitation Learning for Robot Manipulation TasksSimon Stepputtis, Joseph Campbell, Mariano J. Phielipp, Stefan Lee et al.NeurIPS 2020 · 258 citations
- Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic ManipulationXiaoqi Li, Jingyun Xu, Mingxu Zhang, Jiaming Liu et al.CVPR 2025
- RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory SketchesJiayuan Gu, Sean Kirmani, Paul Wohlhart, Yao Lu et al.ICLR 2024 · 135 citations
- Demo2Code: From Summarizing Demonstrations to Synthesizing Code via Extended Chain-of-ThoughtYuki Wang, Gonzalo Gonzalez-Pumariega, Yash Sharma, Sanjiban ChoudhuryNeurIPS 2023 · 67 citations
- Ask Your Humans: Using Human Instructions to Improve Generalization in Reinforcement LearningValerie Chen, Abhinav Gupta, Kenneth MarinoICLR 2021 · 6 citations
