Generate Subgoal Images Before Act: Unlocking the Chain-of-Thought Reasoning in Diffusion Model for Robot Manipulation with Multimodal Prompts
Fei Ni, Jianye Hao, Shiguang Wu, Longxin Kou, Jiashun Liu, Yan Zheng, Bin Wang, Yuzheng Zhuang
摘要
Robotics agents often struggle to understand and follow the multi-modal prompts in complex manipulation scenes which are challenging to be sufficiently and accurately described by text alone. Moreover, for long-horizon manipulation tasks, the deviation from general instruction tends to accumulate if lack of intermediate guidance from high-level subgoals. For this, we consider can we generate subgoal images before act to enhance the instruction following in long-horizon manipulation with multi-modal prompts? Inspired by the great success of diffusion model in image generation tasks, we propose a novel hierarchical framework named as CoTDiffusion that incorporates diffusion model as a high-level planner to convert the general and multimodal prompts into coherent visual subgoal plans, which further guide the low-level policy model before action execution. We design a semantic alignment module that can anchor the progress of generated keyframes along a coherent generation chain, unlocking the chain-of-thought reasoning ability of diffusion model. Additionally, we propose bi-directional generation and frame concat mechanism to further enhance the fidelity of generated subgoal images and the accuracy of instruction following. The experiments cover various robotics manipulation scenarios including visual reasoning, visual rearrange, and visual constraints. CoTDiffusion achieves outstanding performance gain compared to the baselines without explicit subgoal generation, which proves that a subgoal image is worth a thousand words of instruction. The details and visualizations are available at https://cotdiffusion.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- PERIA: Perceive, Reason, Imagine, Act via Holistic Language and Vision Planning for ManipulationFei Ni, Jianye Hao, Shiguang Wu, Longxin Kou 等NeurIPS 2024 · 被引用 13 次
- SheetAgent: Towards a Generalist Agent for Spreadsheet Reasoning and Manipulation via Large Language ModelsYibin Chen, Yifu Yuan, Zeyu Zhang, Yan Zheng 等WWW 2025 · 被引用 13 次
- RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation SkillsChunru Lin, Haotian Yuan, Yian Wang, Xiaowen Qiu 等NeurIPS 2025 · 被引用 10 次
- Target-Aware Video Diffusion ModelsTaeksoo Kim, Hanbyul JooICLR 2026 · 被引用 7 次
- RoboAnnotatorX: A Comprehensive and Universal Annotation Framework for Accurate Understanding of Long-Horizon Robot DemonstrationLongxin Kou, Fei Ni, Yan Zheng, Peilong Han 等ICCV 2025 · 被引用 6 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
相关 Paper
- Instruction-Based Image Editing with Planning, Reasoning, and GenerationLiya Ji, Chenyang Qi, Qifeng ChenICCV 2025 · 被引用 3 次
- CoT-Edit: Let CoT Guide Instruction Video EditingSen Liang, Fengbin Guan, Youliang Zhang, Xin Li 等CVPR 2026 · 被引用 5 次
- CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-stepZheyuan Liu, Munan Ning, Qihui Zhang, Shuo Yang 等NeurIPS 2025 · 被引用 9 次
- SkillDiffuser: Interpretable Hierarchical Planning via Skill Abstractions in Diffusion-Based Task ExecutionZhixuan Liang, Yao Mu, Hengbo Ma, Masayoshi Tomizuka 等CVPR 2024
- HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought ReasoningQuanxin Shou, Fangqi Zhu, Shuang Chen, Puxin Yan 等ICML 2026 · 被引用 7 次
