EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
Siyuan Huang, Liliang Chen, Pengfei Zhou, Shengcong Chen, Yue Liao, Zhengkai Jiang, Yue Hu, Peng Gao, Hongsheng Li, Maoqing Yao, Guanghui Ren
Abstract
We introduce ENERVERSE, a generative robotics foundation model that constructs and interprets embodied spaces. ENERVERSE employs a chunk-wise autoregressive video diffusion framework to predict future embodied spaces from instructions, enhanced by a sparse context memory for long-term reasoning. To model the 3D robotics world, we adopt a multi-view video representation, providing rich perspectives to address challenges like motion ambiguity and 3D grounding. Additionally, ENERVERSE-D, a data engine pipeline combining generative modeling with 4D Gaussian Splatting, forms a self-reinforcing data loop to reduce the sim-to-real gap. Leveraging these innovations, ENERVERSE translates 4D world representations into physical actions via a policy head (ENERVERSE-A), achieving state-of-the-art performance in both simulation and real-world tasks. For efficiency, ENERVERSE-A reuses features from the first denoising step and predicts action chunks, achieving about 280 ms per 8-step action chunk on a single RTX 4090. Further video demos, dataset samples could be found in our project page. * † indicates project leader. ‡ indicates corresponding author. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
as a 'chunk', and the model repeatedly predicts the next chunk to incrementally expand the space. Additionally, to prevent model collapse and enhance the action planning capabilities, we design a sparse context memory mechanism during training. Instead of relying on consecutive memory, this mechanism preserves essential prior content throughout the generation process in a non-redundant manner, theoretically allowing infinite-length sequence generation. While this design achieves stable 2D embodied video generation, it remains insufficient for 3D understanding.
Reonstruction with Obs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Geometry-aware 4D Video Generation for Robot ManipulationZeyi Liu, Shuang Li, Eric Cousineau, Siyuan Feng et al.ICLR 2026 · 28 citations
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action ModelJohn Won, Kyungmin Lee, Huiwon Jang, Dongyoung Kim et al.ICML 2026 · 22 citations
- ORV: 4D Occupancy-centric Robot Video GenerationXiuyu Yang, Bohan Li, Shaocong Xu, Nan Wang et al.CVPR 2026 · 19 citations
- Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level CompositionJiahang Cao, Yize Huang, Hanzhong Guo, Qiang Zhang et al.ICLR 2026 · 14 citations
- Co-Evolving Latent Action World ModelsYucen Wang, Fengming Zhang, De-Chuan Zhan, Li Zhao et al.ICML 2026 · 12 citations
Builds on24
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- Genie Envisioner: A Unified World Foundation Platform for Robotic ManipulationYue Liao, Pengfei Zhou, Siyuan Huang, Donglin Yang et al.ICLR 2026 · 136 citations
- FutureGS: Structured Gaussian Fields for Future-Aware Dynamic Scene ModelingMingyang Ding, Zhan Wang, Jiachen Wang, Tingting Han et al.ACM MM 2025 · 2 citations
- GWM: Towards Scalable Gaussian World Models for Robotic ManipulationGuanxing Lu, Baoxiong Jia, Puhao Li, Yixin Chen et al.ICCV 2025 · 1 citation
- Structured 4D Latent Predictive Model for Robot PlanningZhiyi Li, Peilin Wu, Xiaoshen Han, Ruojin Cai et al.ICML 2026
- ManipDreamer3D: Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D TrajectoryYing Li, Xiaobao Wei, Xiaowei Chi, Yuming Li et al.AAAI 2026
