Structured 4D Latent Predictive Model for Robot Planning
Zhiyi Li, Peilin Wu, Xiaoshen Han, Ruojin Cai, Yilun Du
Abstract
Video predictive models are emerging as a powerful paradigm in robotics, offering a promising path toward task generalization, long-horizon planning, and flexible decision-making. However, prevailing approaches often operate on 2D video sequences, inherently lacking the 3D geometric understanding necessary for precise spatial reasoning and physical consistency. We introduce a Structured 4D Latent Predictive Model , which predicts the evolution of a scene’s 3D structure in a structured latent space conditioned on observations and textual instructions. Our representation encodes the scene holistically and can be decoded into diverse 3D formats, enabling a more complete and 3D consistent scene understanding. This structured 4D latent predictive model serves as a planner, generating future scenes that are translated into executable actions by a goal-conditioned inverse dynamics module. Experiments demonstrate that our model generates futures with strong visual quality, substantially better 3D consistency and multi-view coherence compared to state-of-the-art video-based planners. Consequently, our full planning pipeline achieves superior performance on complex manipulation tasks, exhibits robust generalization to novel visual conditions, and proves effective on real-world robotic platforms. Our website is available at https://structured-4d-model.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f193cffd-cb36-4e7f-ae18-2458d40c1c85Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz et al.NeurIPS 2024 · 751 citations
Related papers
- Geometry-aware 4D Video Generation for Robot ManipulationZeyi Liu, Shuang Li, Eric Cousineau, Siyuan Feng et al.ICLR 2026 · 28 citations
- AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D ScenesYu Li, Menghan Xia, Gongye Liu, Jianhong Bai et al.ICLR 2026 · 3 citations
- 4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion ModelsHeng Yu, Chaoyang Wang, Peiye Zhuang, Willi Menapace et al.NeurIPS 2024 · 76 citations
- Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-MotionNils Morbitzer, Jonathan Evers, Artem Savkin, Thomas Stauner et al.ICML 2026
- MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic ManipulationJiaxu Wang, JIANG Yicheng, Tianlun HE, Jingkai SUN et al.ICML 2026 · 8 citations
