Lune

NeurIPS2025Top-tier venue

EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation

Siyuan Huang, Liliang Chen, Pengfei Zhou, Shengcong Chen, Yue Liao, Zhengkai Jiang, Yue Hu, Peng Gao, Hongsheng Li, Maoqing Yao, Guanghui Ren

2025Year
66Citations
10Top-tier citations

Abstract

We introduce ENERVERSE, a generative robotics foundation model that constructs and interprets embodied spaces. ENERVERSE employs a chunk-wise autoregressive video diffusion framework to predict future embodied spaces from instructions, enhanced by a sparse context memory for long-term reasoning. To model the 3D robotics world, we adopt a multi-view video representation, providing rich perspectives to address challenges like motion ambiguity and 3D grounding. Additionally, ENERVERSE-D, a data engine pipeline combining generative modeling with 4D Gaussian Splatting, forms a self-reinforcing data loop to reduce the sim-to-real gap. Leveraging these innovations, ENERVERSE translates 4D world representations into physical actions via a policy head (ENERVERSE-A), achieving state-of-the-art performance in both simulation and real-world tasks. For efficiency, ENERVERSE-A reuses features from the first denoising step and predicts action chunks, achieving about 280 ms per 8-step action chunk on a single RTX 4090. Further video demos, dataset samples could be found in our project page. * † indicates project leader. ‡ indicates corresponding author. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

as a 'chunk', and the model repeatedly predicts the next chunk to incrementally expand the space. Additionally, to prevent model collapse and enhance the action planning capabilities, we design a sparse context memory mechanism during training. Instead of relying on consecutive memory, this mechanism preserves essential prior content throughout the generation process in a non-redundant manner, theoretically allowing infinite-length sequence generation. While this design achieves stable 2D embodied video generation, it remains insufficient for 3D understanding.

Reonstruction with Obs.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers10

Ask how each one uses it

Builds on24

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines