ManipDreamer3D: Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
Ying Li, Xiaobao Wei, Xiaowei Chi, Yuming Li, Zhongyu Zhao, Hao Wang, Ningning Ma, Ming Lu, Sirui Han
Abstract
Data scarcity continues to be a critical bottleneck in the field of robotic manipulation, limiting the ability to train robust and generalizable models. While diffusion models provide a promising approach to synthesizing realistic robotic manipulation videos, their effectiveness hinges on the availability of precise and reasonable control instructions. Current methods primarily rely on 2D trajectories as instruction prompts, which inherently face issues with 3D spatial ambiguity. In this work, we present a novel framework named ManipDreamer3Dfor generating plausible 3D-aware robotic manipulation videos from the input image and the text instruction. Our method combines 3D trajectory planning with a reconstructed 3D occupancy map created from a third-person perspective, along with a novel trajectory-to-video diffusion model. Specifically, ManipDreamer3D first reconstructs the 3D occupancy representation from the input image and then computes an optimized 3D end-effector trajectory, minimizing path length, avoiding collisions and retiming. Next, we employ a latent editing technique to create video sequences from the initial image latent, text instruction and the optimized 3D trajectory. This process conditions our specially trained trajectory-to-video diffusion model to produce robotic pick-and-place videos. Our method significantly reduces human intervention requirements by autonomously planing plausible 3D trajectories. Experimental results demonstrate its superior visual quality and precision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a8cfe84-abe7-4fd9-af4a-12931e7755a0Cited by top-tier papers5
- OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic CameraHao Shi, Ze Wang, Shangwei Guo, Mengfei Duan et al.CVPR 2026 · 11 citations
- ParkGaussian: Surround-view 3D Gaussian Splatting for Autonomous ParkingXiaobao Wei, Zhangjie Ye, Yuxiang Gu, Zunjie Zhu et al.CVPR 2026 · 8 citations
- RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic ManipulationSixu Lin, Junliang Chen, Huaiyuan Xu, Zhuohao Li et al.ICML 2026 · 3 citations
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and CompressionYilong Chen, Xiang Bai, Zhibin Wang, Chengyu Bai et al.AAAI 2026 · 1 citation
- Predicting What Matters: Robust Generalist Robot Policy Learning via Future Semantic MaskYunfan Lou, Xiaowei Chi, Xiaojie Zhang, Zezhong Qian et al.ICML 2026
Builds on9
- MotionCtrl: A Unified and Flexible Motion Controller for Video GenerationZhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li et al.SIGGRAPH 2024 · 123 citations
- Learning Video Generation for Robotic Manipulation with Collaborative Trajectory ControlXiao Fu, Xintao Wang, Xian Liu, Jianhong Bai et al.ICLR 2026 · 37 citations
- MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language NavigationJiaqi Chen, Bingqian Lin, Ran Xu, Zhenhua Chai et al.ACL 2024 · 29 citations
- Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation ControlZekai Gu, Rui Yan, Jiahao Lu, Peng Li et al.SIGGRAPH 2025 · 21 citations
- FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token PruningJiajun Cao, Qizhe Zhang, Peidong Jia, Xuhui Zhao et al.AAAI 2026 · 18 citations
Related papers
- General Articulated Objects Manipulation in Real Images via Part-Aware Diffusion ProcessZhou Fang, Yong-Lu Li, Lixin Yang, Cewu LuNeurIPS 2024
- Target-Aware Video Diffusion ModelsTaeksoo Kim, Hanbyul JooICLR 2026 · 7 citations
- Towards Physical Understanding in Video Generation: A 3D Point Regularization ApproachYunuo Chen, Junli Cao, Vidit Goel, Sergei Korolev et al.NeurIPS 2025 · 9 citations
- HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video SynthesisMingjin Chen, Junhao Chen, Zhaoxin Fan, Yujian Lee et al.CVPR 2026 · 13 citations
- HOIDiffusion: Generating Realistic 3D Hand-Object Interaction DataMengqi Zhang, Yang Fu, Zheng Ding, Sifei Liu et al.CVPR 2024 · 18 citations
