Pretrained Language Models as Visual Planners for Human Assistance
Dhruvesh Patel, Hamid Eghbalzadeh, Nitin Kamra, Michael Louis Iuzzolino, Unnat Jain, Ruta Desai
Abstract
In our pursuit of advancing multi-modal AI assistants capable of guiding users to achieve complex multi-step goals, we propose the task of ‘Visual Planning for Assistance (VPA)’. Given a succinct natural language goal, e.g., "make a shelf", and a video of the user’s progress so far, the aim of VPA is to devise a plan, i.e. a sequence of actions such as "sand shelf", "paint shelf", etc. to realize the specified goal. This requires assessing the user’s progress from the (untrimmed) video, and relating it to the requirements of natural language goal, i.e. which actions to select and in what order? Consequently, this requires handling long video history and arbitrarily complex action dependencies. To address these challenges, we decompose VPA into video action segmentation and forecasting. Importantly, we experiment by formulating the forecasting step as a multimodal sequence modeling problem, allowing us to leverage the strength of pre-trained LMs (as the sequence model). This novel approach, which we call Visual Language Model based Planner (VLaMP), outperforms baselines across a suite of metrics that gauge the quality of the generated plans. Furthermore, through comprehensive ablations, we also isolate the value of each component – language pre-training, visual observations, and goal information. We have open-sourced all the data, model checkpoints, and training code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6b4b60c-be33-4050-a2e1-567aacbd3a5fCited by top-tier papers9
- Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied AgentsWenlong Huang, Fei Xia, Dhruv Shah, Danny Driess et al.NeurIPS 2023 · 102 citations
- Robotic Manipulation by Imitating Generated Videos Without Physical DemonstrationsShivansh Patel, Shraddhaa Mohan, Hanlin Mai, Unnat Jain et al.ICLR 2026 · 50 citations
- VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained ActionsGuangyan Chen, Meiling Wang, Te Cui, Yao Mu et al.NeurIPS 2024 · 24 citations
- VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory BridgesYuxuan Wang, Yiqi Song, Cihang Xie, Yang Liu et al.ICCV 2025 · 7 citations
- GeoWorld: Geometric World ModelsZeyu Zhang, Danning Li, Ian Reid, Richard HartleyCVPR 2026 · 6 citations
Builds on34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- Video Language PlanningYilun Du, Sherry Yang, Pete Florence, Fei Xia et al.ICLR 2024 · 161 citations
- LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural PlanningShibo Sun, Xue Li, Donglin Di, Mingjie Wei et al.ACM MM 2025 · 4 citations
- Show and Guide: Instructional-Plan Grounded Vision and Language ModelDiogo Glória-Silva, David Semedo, João MagalhãesEMNLP 2024
- LAMP: Language-Assisted Motion Planning for Controllable Video GenerationMuhammed Burak Kizil, Enes Şanlı, Niloy J. Mitra, Erkut Erdem et al.CVPR 2026 · 4 citations
- VideoVLA: Video Generators Can Be Generalizable Robot ManipulatorsYichao Shen, Fangyun Wei, Zhiying Du, Yaobo Liang et al.NeurIPS 2025 · 73 citations
