Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
Yue Liao, Pengfei Zhou, Siyuan Huang, Donglin Yang, Shengcong Chen, Yuxin Jiang, Yue Hu, Si Liu, Jianlan Luo, Liliang Chen, Shuicheng Yan, Maoqing Yao, Guanghui Ren
Abstract
We introduce Genie Envisioner (GE), a unified world foundation platform for robotic manipulation that integrates policy learning, evaluation, and simulation within a single video-generative framework. At its core, GE-Base is a large-scale, instruction-conditioned video diffusion model that captures the spatial, temporal, and semantic dynamics of real-world robotic interactions in a structured latent space. Built upon this foundation, GE-Act maps latent representations to executable action trajectories through a lightweight, flow-matching decoder, enabling precise and generalizable policy inference across diverse embodiments with minimal supervision. To support scalable evaluation and training, GE-Sim serves as an action-conditioned neural simulator, producing high-fidelity rollouts for closed-loop policy development. The platform is further equipped with EWMBench, a standardized benchmark suite measuring visual fidelity, physical consistency, and instruction-action alignment. Together, these components establish Genie Envisioner as a scalable and practical foundation for instruction-driven, general-purpose embodied intelligence. All code, models, and benchmarks will be released publicly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d49f996-ea84-4475-98b3-866a11ef9e4dCited by top-tier papers15
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and PlanningMoo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin et al.ICLR 2026 · 325 citations
- Ctrl-World: A Controllable Generative World Model for Robot ManipulationYanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea FinnICLR 2026 · 163 citations
- DriveLaW: Unifying Planning and Video Generation in a Latent Driving WorldTianze Xia, Yongkang Li, Lijun Zhou, Jingfeng Yao et al.CVPR 2026 · 58 citations
- ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action ModelsLinqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong et al.CVPR 2026 · 43 citations
- World Guidance: World Modeling in Condition Space for Action GenerationYue Su, Sijin Chen, Haixin Shi, Mingyu Liu et al.ICML 2026 · 26 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder et al.ICML 2024 · 513 citations
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai et al.ICML 2026 · 394 citations
Related papers
- GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following ManipulationNing Gao, Yilun Chen, Shuai Yang, Xinyi Chen et al.CVPR 2025
- DreamDojo: A Real-Time Robot World Model from Large-Scale Human VideosShenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik et al.ICML 2026 · 96 citations
- VITA: Vision-to-Action Flow Matching PolicyDechen Gao, BOQI ZHAO, Andrew Lee, Ian Chuang et al.ICLR 2026 · 27 citations
- Rethinking Video Generation Model for the Embodied WorldYufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li et al.ICML 2026 · 24 citations
- VideoVLA: Video Generators Can Be Generalizable Robot ManipulatorsYichao Shen, Fangyun Wei, Zhiying Du, Yaobo Liang et al.NeurIPS 2025 · 73 citations
