MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation
Jiaxu Wang, JIANG Yicheng, Tianlun HE, Jingkai SUN, Qiang Zhang, Jiahang Cao, Zesen Gan, Mingyuan Sun, Qiming Shao, Xiangyu Yue
摘要
World-model-based imagine-then-act becomes a promising paradigm for robotic manipulation, yet existing approaches typically support either purely image-based forecasting or reasoning over partial 3D geometry, limiting their ability to predict complete 4D scene dynamics. To solve this, this work explores a novel embodied 4D world model that enables geometrically consistent, arbitrary-view RGBD generation: given only a single-view RGBD observation as input, the model “imagines” the remaining viewpoints, which can then be back-projected and fused to assemble a more complete 3D structure across time. To efficiently learn the multi-view, cross-modality generation, we explicitly design cross-view and cross-modality feature fusion that jointly encourage consistency between RGB and depth and enforce geometric alignment across views. Beyond prediction, converting generated futures into actions is often handled by inverse dynamics, which is ill-posed because multiple actions can explain the same transition. We address this with a test-time action optimization strategy that backpropagates through the generative model to infer a trajectory-level latent best matching the predicted future, and a residual inverse dynamics model that turns this trajectory prior into accurate executable actions. Extensive experiments on the three datasets and platforms demonstrate strong performance on both 4D scene generation and downstream manipulation, and ablations provide practical insights into the key design choices. Project page is available at https://mercerai.github.io/MVISTA-4D/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai 等NeurIPS 2023 · 被引用 742 次
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen 等ICLR 2026 · 被引用 720 次
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai 等ICML 2026 · 被引用 394 次
相关 Paper
- Geometry-aware 4D Video Generation for Robot ManipulationZeyi Liu, Shuang Li, Eric Cousineau, Siyuan Feng 等ICLR 2026 · 被引用 28 次
- Learning 4D Embodied World ModelsHaoyu Zhen, Qiao Sun, Hongxin Zhang, Junyan Li 等ICCV 2025 · 被引用 5 次
- WorldReel: 4D Video Generation with Consistent Geometry and Motion ModelingShaoheng Fang, Hanwen Jiang, Yunpeng Bai, Niloy J. Mitra 等CVPR 2026 · 被引用 3 次
- Structured 4D Latent Predictive Model for Robot PlanningZhiyi Li, Peilin Wu, Xiaoshen Han, Ruojin Cai 等ICML 2026
- Restage4D: Reanimating Deformable 3D Reconstruction from a Single VideoJixuan He, Chieh Hubert Lin, Lu Qi, Ming-Hsuan YangNeurIPS 2025 · 被引用 2 次
