ManipDreamer3D: Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
Ying Li, Xiaobao Wei, Xiaowei Chi, Yuming Li, Zhongyu Zhao, Hao Wang, Ningning Ma, Ming Lu, Sirui Han
摘要
Data scarcity continues to be a critical bottleneck in the field of robotic manipulation, limiting the ability to train robust and generalizable models. While diffusion models provide a promising approach to synthesizing realistic robotic manipulation videos, their effectiveness hinges on the availability of precise and reasonable control instructions. Current methods primarily rely on 2D trajectories as instruction prompts, which inherently face issues with 3D spatial ambiguity. In this work, we present a novel framework named ManipDreamer3Dfor generating plausible 3D-aware robotic manipulation videos from the input image and the text instruction. Our method combines 3D trajectory planning with a reconstructed 3D occupancy map created from a third-person perspective, along with a novel trajectory-to-video diffusion model. Specifically, ManipDreamer3D first reconstructs the 3D occupancy representation from the input image and then computes an optimized 3D end-effector trajectory, minimizing path length, avoiding collisions and retiming. Next, we employ a latent editing technique to create video sequences from the initial image latent, text instruction and the optimized 3D trajectory. This process conditions our specially trained trajectory-to-video diffusion model to produce robotic pick-and-place videos. Our method significantly reduces human intervention requirements by autonomously planing plausible 3D trajectories. Experimental results demonstrate its superior visual quality and precision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic CameraHao Shi, Ze Wang, Shangwei Guo, Mengfei Duan 等CVPR 2026 · 被引用 11 次
- ParkGaussian: Surround-view 3D Gaussian Splatting for Autonomous ParkingXiaobao Wei, Zhangjie Ye, Yuxiang Gu, Zunjie Zhu 等CVPR 2026 · 被引用 8 次
- RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic ManipulationSixu Lin, Junliang Chen, Huaiyuan Xu, Zhuohao Li 等ICML 2026 · 被引用 3 次
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and CompressionYilong Chen, Xiang Bai, Zhibin Wang, Chengyu Bai 等AAAI 2026 · 被引用 1 次
- Predicting What Matters: Robust Generalist Robot Policy Learning via Future Semantic MaskYunfan Lou, Xiaowei Chi, Xiaojie Zhang, Zezhong Qian 等ICML 2026
它引用的顶会 Paper9
- MotionCtrl: A Unified and Flexible Motion Controller for Video GenerationZhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li 等SIGGRAPH 2024 · 被引用 123 次
- Learning Video Generation for Robotic Manipulation with Collaborative Trajectory ControlXiao Fu, Xintao Wang, Xian Liu, Jianhong Bai 等ICLR 2026 · 被引用 37 次
- MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language NavigationJiaqi Chen, Bingqian Lin, Ran Xu, Zhenhua Chai 等ACL 2024 · 被引用 29 次
- Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation ControlZekai Gu, Rui Yan, Jiahao Lu, Peng Li 等SIGGRAPH 2025 · 被引用 21 次
- FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token PruningJiajun Cao, Qizhe Zhang, Peidong Jia, Xuhui Zhao 等AAAI 2026 · 被引用 18 次
相关 Paper
- General Articulated Objects Manipulation in Real Images via Part-Aware Diffusion ProcessZhou Fang, Yong-Lu Li, Lixin Yang, Cewu LuNeurIPS 2024
- Target-Aware Video Diffusion ModelsTaeksoo Kim, Hanbyul JooICLR 2026 · 被引用 7 次
- Towards Physical Understanding in Video Generation: A 3D Point Regularization ApproachYunuo Chen, Junli Cao, Vidit Goel, Sergei Korolev 等NeurIPS 2025 · 被引用 9 次
- HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video SynthesisMingjin Chen, Junhao Chen, Zhaoxin Fan, Yujian Lee 等CVPR 2026 · 被引用 13 次
- HOIDiffusion: Generating Realistic 3D Hand-Object Interaction DataMengqi Zhang, Yang Fu, Zheng Ding, Sifei Liu 等CVPR 2024 · 被引用 18 次
