Autoscape: Geometry-Consistent Long-Horizon Scene Generation
Jiacheng Chen, Ziyu Jiang, Mingfu Liang, Bingbing Zhuang, Jong-Chyi Su, Sparsh Garg, Ying Wu, Manmohan Chandraker
Abstract
This paper proposes AutoScape, a long-horizon driving scene generation framework. At its core is a novel RGB-D diffusion model that iteratively generates sparse, geometrically consistent keyframes, serving as reliable anchors for the scene's appearance and geometry. To maintain long-range geometric consistency, the model 1) jointly handles image and depth in a shared latent space, 2) explicitly conditions on the existing scene geometry (i.e., rendered point clouds) from previously generated keyframes, and 3) steers the sampling process with a warp-consistent guidance. Given high-quality RGB-D keyframes, a video diffusion model then interpolates between them to produce dense and coherent video frames. AutoScape generates realistic and geometrically consistent driving videos of over 20 seconds, improving the long-horizon FID and FVD scores over the prior state-of-the-art by 48.6% and 43.0%, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4ae92f5-1924-450e-922d-4064a6f83f22Cited by top-tier papers2
- HorizonForge: Driving Scene Editing with Any Trajectories and Any VehiclesYifan Wang, Francesco Pittaluga, Zaid Tasneem, Chenyu You et al.CVPR 2026 · 3 citations
- ScenDi: 3D-to-2D Scene Diffusion Cascades for Urban GenerationHanlei Guo, Jiahao Shao, Xinya Chen, Xiyang Tan et al.CVPR 2026 · 1 citation
Builds on56
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
Related papers
- X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible ControllabilityYu Yang, Alan Liang, Jianbiao Mei, Yukai Ma et al.NeurIPS 2025 · 22 citations
- DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene GenerationJiazhe Guo, Yikang Ding, Xiwu Chen, Shuo Chen et al.ICCV 2025 · 5 citations
- Grounded Latents for Entity-Centric 4D Scene GenerationJinhyung Park, Navyata Sanghvi, Erica Weng, Shawn Hunt et al.CVPR 2026
- DriveScape: High-Resolution Driving Video Generation by Multi-View Feature FusionWei Wu, Xi Guo, Weixuan Tang, Tingxuan Huang et al.CVPR 2025
- MultiDiff: Consistent Novel View Synthesis from a Single ImageNorman Müller, Katja Schwarz, Barbara Rössle, Lorenzo Porzi et al.CVPR 2024 · 14 citations
