Zero-Shot Robotic Manipulation with Pre-Trained Image-Editing Diffusion Models
Kevin Black, Mitsuhiko Nakamoto, Pranav Atreya, Homer Rich Walke, Chelsea Finn, Aviral Kumar, Sergey Levine
摘要
If generalist robots are to operate in truly unstructured environments, they need to be able to recognize and reason about novel objects and scenarios. Such objects and scenarios might not be present in the robot's own training data. We propose SuSIE, a method that leverages an image-editing diffusion model to act as a highlevel planner by proposing intermediate subgoals that a low-level controller can accomplish. Specifically, we finetune InstructPix2Pix on video data, consisting of both human videos and robot rollouts, such that it outputs hypothetical future "subgoal" observations given the robot's current observation and a language command. We also use the robot data to train a low-level goal-conditioned policy to act as the aforementioned low-level controller. We find that the high-level subgoal predictions can utilize Internet-scale pretraining and visual understanding to guide the low-level goal-conditioned policy, achieving significantly better generalization and precision than conventional language-conditioned policies. We achieve state-of-the-art results on the CALVIN benchmark, and also demonstrate robust generalization on real-world manipulation tasks, beating strong baselines that have access to privileged information or that utilize orders of magnitude more compute and training data. The project website can be found at http://rail-berkeley.github.io/susie .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper95
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta 等NeurIPS 2024 · 被引用 403 次
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelJinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu 等ICLR 2026 · 被引用 335 次
- 3D-VLA: A 3D Vision-Language-Action Generative World ModelHaoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang 等ICML 2024 · 被引用 303 次
- Motus: A Unified Latent Action World ModelHongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang 等CVPR 2026 · 被引用 271 次
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang 等NeurIPS 2025 · 被引用 244 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
相关 Paper
- Language Control Diffusion: Efficiently Scaling through Space, Time, and TasksEdwin Zhang, Yujie Lu, Shinda Huang, William Yang Wang 等ICLR 2024 · 被引用 34 次
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationHongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen 等ICLR 2024 · 被引用 309 次
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsYucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen 等ICML 2025
- Generate Subgoal Images Before Act: Unlocking the Chain-of-Thought Reasoning in Diffusion Model for Robot Manipulation with Multimodal PromptsFei Ni, Jianye Hao, Shiguang Wu, Longxin Kou 等CVPR 2024 · 被引用 4 次
- Pixel Motion Diffusion is What We Need for Robot ControlE-Ro Nguyen, Yichi Zhang, Kanchana Ranasinghe, Xiang Li 等CVPR 2026 · 被引用 10 次
