GenEx: Generating an Explorable World
Taiming Lu, Tianmin Shu, Alan L. Yuille, Daniel Khashabi, Jieneng Chen
Abstract
Understanding, navigating, and exploring the 3D physical real world has long been a central challenge in the development of artificial intelligence. In this work, we take a step toward this goal by introducing GenEx, a system capable of planning complex embodied world exploration, guided by its generative imagination that forms expectations about the surrounding environments. GenEx generates high-quality, continuous 360-degree virtual environments, achieving high loop consistency and active 3D mapping over extended trajectories. Leveraging generative imagination, GPT-assisted agents can undertake complex embodied tasks, including goal-agnostic exploration and goal-driven navigation. Agents utilize imagined observations to update their beliefs, simulate potential outcomes, and enhance their decision-making. Training on the synthetic urban dataset GenEx-DB and evaluation on GenEx-EQA demonstrate that our approach significantly improves agents' planning capabilities, providing a transformative platform toward intelligent, imaginative embodied exploration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78fae8e7-a6b8-45bb-abab-705cde659fb7Cited by top-tier papers7
- PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware MechanismsYifei Xia, Shuchen Weng, Siqi Yang, Jingqi Liu et al.NeurIPS 2025 · 24 citations
- DA2: Depth Anything in Any DirectionHaodong Li, Wangguandong Zheng, Jing He, Yuhao Liu et al.ICLR 2026 · 23 citations
- TV2TV: A Unified Framework for Interleaved Language and Video GenerationXiaochuang Han, Youssef Emad, Melissa Hall, John Nguyen et al.CVPR 2026 · 3 citations
- Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video DiffusionTing-Hsuan Chen, Ying-Huan Chen, Tao Tu, Jie-Ying Lee et al.CVPR 2026 · 2 citations
- 360Explorer: Exploring 4D Controllable World in Panoramic VideosXinhua Cheng, Haiyang Zhou, Wangbo Yu, Tanghui Jia et al.AAAI 2026
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
Related papers
- Holodeck: Language Guided Generation of 3D Embodied AI EnvironmentsYue Yang, Fan-Yun Sun, Luca Weihs, Eli VanderBilt et al.CVPR 2024 · 47 citations
- EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of ThoughtYao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang et al.NeurIPS 2023 · 453 citations
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- Towards Learning a Generalist Model for Embodied NavigationDuo Zheng, Shijia Huang, Lin Zhao, Yiwu Zhong et al.CVPR 2024 · 37 citations
