Grounding Language Plans in Demonstrations Through Counterfactual Perturbations
Yanwei Wang, Tsun-Hsuan Wang, Jiayuan Mao, Michael Hagenow, Julie Shah
Abstract
Grounding the common-sense reasoning of Large Language Models (LLMs) in physical domains remains a pivotal yet unsolved problem for embodied AI. Whereas prior works have focused on leveraging LLMs directly for planning in symbolic spaces, this work uses LLMs to guide the search of task structures and constraints implicit in multi-step demonstrations. Specifically, we borrow from manipulation planning literature the concept of mode families, which group robot configurations by specific motion constraints, to serve as an abstraction layer between the high-level language representations of an LLM and the low-level physical trajectories of a robot. By replaying a few human demonstrations with synthetic perturbations, we generate coverage over the demonstrations' state space with additional successful executions as well as counterfactuals that fail the task. Our explanation-based learning framework trains an end-to-end differentiable neural network to predict successful trajectories from failures and as a by-product learns classifiers that ground low-level states and images in mode families without dense labeling. The learned grounding classifiers can further be used to translate language plans into reactive policies in the physical domain in an interpretable manner. We show our approach improves the interpretability and reactivity of imitation learning through 2D navigation and simulated and real robot manipulation tasks. Website: https://yanweiw.github.io/glide/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Pre-Trained Language Models for Interactive Decision-MakingShuang Li, Xavier Puig, Chris Paxton, Yilun Du et al.NeurIPS 2022 · 341 citations
- Skill Induction and Planning with Latent LanguagePratyusha Sharma, Antonio Torralba, Jacob AndreasACL 2022 · 127 citations
- Learning Rational Subgoals from Demonstrations and InstructionsZhezheng Luo, Jiayuan Mao, Jiajun Wu, Tomás Lozano-Pérez et al.AAAI 2023 · 5 citations
Related papers
- ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic ManipulationXiaoqi Li, Mingxu Zhang, Yiran Geng, Haoran Geng et al.CVPR 2024
- Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied AgentsWenlong Huang, Fei Xia, Dhruv Shah, Danny Driess et al.NeurIPS 2023 · 102 citations
- Instruction-Augmented Long-Horizon Planning: Embedding Grounding Mechanisms in Embodied Mobile ManipulationFangyuan Wang, Shipeng Lyu, Peng Zhou, Anqing Duan et al.AAAI 2025 · 9 citations
- Hypothesis, Verification, and Induction: Grounding Large Language Models with Self-Driven Skill LearningShaohui Peng, Xing Hu, Qi Yi, Rui Zhang et al.AAAI 2024 · 4 citations
- Towards Reliable Code-as-Policies: A Neuro-Symbolic Framework for Embodied Task PlanningSanghyun Ahn, Wonje Choi, Junyong Lee, Jinwoo Park et al.NeurIPS 2025 · 14 citations
