Generate Subgoal Images Before Act: Unlocking the Chain-of-Thought Reasoning in Diffusion Model for Robot Manipulation with Multimodal Prompts
Fei Ni, Jianye Hao, Shiguang Wu, Longxin Kou, Jiashun Liu, Yan Zheng, Bin Wang, Yuzheng Zhuang
Abstract
Robotics agents often struggle to understand and follow the multi-modal prompts in complex manipulation scenes which are challenging to be sufficiently and accurately described by text alone. Moreover, for long-horizon manipulation tasks, the deviation from general instruction tends to accumulate if lack of intermediate guidance from high-level subgoals. For this, we consider can we generate subgoal images before act to enhance the instruction following in long-horizon manipulation with multi-modal prompts? Inspired by the great success of diffusion model in image generation tasks, we propose a novel hierarchical framework named as CoTDiffusion that incorporates diffusion model as a high-level planner to convert the general and multimodal prompts into coherent visual subgoal plans, which further guide the low-level policy model before action execution. We design a semantic alignment module that can anchor the progress of generated keyframes along a coherent generation chain, unlocking the chain-of-thought reasoning ability of diffusion model. Additionally, we propose bi-directional generation and frame concat mechanism to further enhance the fidelity of generated subgoal images and the accuracy of instruction following. The experiments cover various robotics manipulation scenarios including visual reasoning, visual rearrange, and visual constraints. CoTDiffusion achieves outstanding performance gain compared to the baselines without explicit subgoal generation, which proves that a subgoal image is worth a thousand words of instruction. The details and visualizations are available at https://cotdiffusion.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b941eb9a-01b9-47d0-939a-ae6fdfdf8e1cCited by top-tier papers15
- PERIA: Perceive, Reason, Imagine, Act via Holistic Language and Vision Planning for ManipulationFei Ni, Jianye Hao, Shiguang Wu, Longxin Kou et al.NeurIPS 2024 · 13 citations
- SheetAgent: Towards a Generalist Agent for Spreadsheet Reasoning and Manipulation via Large Language ModelsYibin Chen, Yifu Yuan, Zeyu Zhang, Yan Zheng et al.WWW 2025 · 13 citations
- RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation SkillsChunru Lin, Haotian Yuan, Yian Wang, Xiaowen Qiu et al.NeurIPS 2025 · 10 citations
- Target-Aware Video Diffusion ModelsTaeksoo Kim, Hanbyul JooICLR 2026 · 7 citations
- RoboAnnotatorX: A Comprehensive and Universal Annotation Framework for Accurate Understanding of Long-Horizon Robot DemonstrationLongxin Kou, Fei Ni, Yan Zheng, Peilong Han et al.ICCV 2025 · 6 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
Related papers
- Instruction-Based Image Editing with Planning, Reasoning, and GenerationLiya Ji, Chenyang Qi, Qifeng ChenICCV 2025 · 3 citations
- CoT-Edit: Let CoT Guide Instruction Video EditingSen Liang, Fengbin Guan, Youliang Zhang, Xin Li et al.CVPR 2026 · 5 citations
- CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-stepZheyuan Liu, Munan Ning, Qihui Zhang, Shuo Yang et al.NeurIPS 2025 · 9 citations
- SkillDiffuser: Interpretable Hierarchical Planning via Skill Abstractions in Diffusion-Based Task ExecutionZhixuan Liang, Yao Mu, Hengbo Ma, Masayoshi Tomizuka et al.CVPR 2024
- HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought ReasoningQuanxin Shou, Fangqi Zhu, Shuang Chen, Puxin Yan et al.ICML 2026 · 7 citations
