Learning 4D Embodied World Models
Haoyu Zhen, Qiao Sun, Hongxin Zhang, Junyan Li, Siyuan Zhou, Yilun Du, Chuang Gan
摘要
This paper presents an effective approach for learning novel 4D embodied world models, TesserAct, which predict the dynamic evolution of 3D scenes over time in response to an embodied agent's actions, providing both spatial and temporal consistency. We propose to learn a 4D world model by training on RGB-DN (RGB, Depth, and Normal) videos. This not only surpasses traditional 2D models by incorporating detailed shape, configuration, and temporal changes into their predictions, but also allows us to effectively learn accurate inverse dynamic models for an embodied agent. Specifically, we first extend existing robotic manipulation video datasets
with depth and normal information leveraging off-the-shelf models. Next, we fine-tune a video generation model on this annotated dataset, which jointly predicts RGB-DN (RGB, Depth, and Normal) for each frame. We then present an algorithm to directly convert generated RGB, Depth, and Normal videos into a high-quality 4D scene of the world.
Our method ensures temporal and spatial coherence in 4D scene predictions from embodied scenarios, enables novel view synthesis for embodied environments, and facilitates policy learning that significantly outperforms those derived from prior video-based world models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- World-In-World: World Models in a Closed-Loop WorldJiahan Zhang, Muqing Jiang, Nanru Dai, Taiming Lu 等ICLR 2026 · 被引用 46 次
- RoboScape: Physics-informed Embodied World ModelYu Shang, Xin Zhang, Yinzhou Tang, Lei Jin 等NeurIPS 2025 · 被引用 43 次
- Learning World Models for Interactive Video GenerationTaiye Chen, Xun Hu, Zihan Ding, Chi JinNeurIPS 2025 · 被引用 37 次
- Learning 3D Persistent Embodied World ModelsSiyuan Zhou, Yilun Du, Yuncong Yang, Lei Han 等NeurIPS 2025 · 被引用 34 次
- Geometry-aware 4D Video Generation for Robot ManipulationZeyi Liu, Shuang Li, Eric Cousineau, Siyuan Feng 等ICLR 2026 · 被引用 28 次
它引用的顶会 Paper35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic ManipulationJiaxu Wang, JIANG Yicheng, Tianlun HE, Jingkai SUN 等ICML 2026 · 被引用 8 次
- Structured 4D Latent Predictive Model for Robot PlanningZhiyi Li, Peilin Wu, Xiaoshen Han, Ruojin Cai 等ICML 2026
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World ModelingKairun Wen, Yuzhi Huang, Runyu Chen, Hui Zheng 等NeurIPS 2025 · 被引用 11 次
- WorldReel: 4D Video Generation with Consistent Geometry and Motion ModelingShaoheng Fang, Hanwen Jiang, Yunpeng Bai, Niloy J. Mitra 等CVPR 2026 · 被引用 3 次
- Understanding Dynamic Scenes in Ego Centric 4D Point CloudsJunsheng Huang, Shengyu Hao, Bocheng Hu, Hongwei Wang 等AAAI 2026 · 被引用 4 次
