StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
Yunzhi Yan, Zhen Xu, Haotong Lin, Haian Jin, Haoyu Guo, Yida Wang, Kun Zhan, Xianpeng Lang, Hujun Bao, Xiaowei Zhou, Sida Peng
Abstract
This paper aims to tackle the problem of photorealistic view synthesis from vehicle sensor data. Recent advancements in neural scene representation have achieved notable success in rendering high-quality autonomous driving scenes, but the performance significantly degrades as the viewpoint deviates from the training trajectory. To mitigate this problem, we introduce StreetCrafter, a novel controllable video diffusion model that utilizes LiDAR point cloud renderings as pixel-level conditions, which fully exploits the generative prior for novel view synthesis, while preserving precise camera control. Moreover, the utilization of pixel-level LiDAR conditions allows us to make accurate pixel-level edits to target scenes. In addition, the generative prior of StreetCrafter can be effectively incorporated into dynamic scene representations to achieve real-time rendering. Experiments on Waymo Open Dataset and PandaSet demonstrate that our model enables flexible control over viewpoint changes, enlarging the view synthesis regions for satisfying rendering, which outperforms existing methods. The code is available at https://zju3dv.github.io/street crafter.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3910aa6-4eeb-485d-827b-1f483c03a1ccCited by top-tier papers13
- X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible ControllabilityYu Yang, Alan Liang, Jianbiao Mei, Yukai Ma et al.NeurIPS 2025 · 22 citations
- From Rays to Projections: Better Inputs for Feed-Forward View SynthesisZirui Wu, Zeren Jiang, Martin R. Oswald, Jie SongCVPR 2026 · 5 citations
- DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene GenerationJiazhe Guo, Yikang Ding, Xiwu Chen, Shuo Chen et al.ICCV 2025 · 5 citations
- Causal-Entity Reflected Egocentric Traffic Accident Video SynthesisLei-Lei Li, Jianwu Fang, Junbin Xiao, Shanmin Pang et al.ICCV 2025 · 4 citations
- UMAMI: Unifying Masked Autoregressive Models and Deterministic Rendering for View SynthesisThanh-Tung Le, Tuan Pham, Tung Nguyen, Deying Kong et al.NeurIPS 2025 · 4 citations
Builds on48
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- LidarPainter: One-Step Away from Any Lidar View to Novel GuidanceYuzhou Ji, Ke Ma, Hong Cai, Anchun Zhang et al.AAAI 2026
- LiDARCrafter: Dynamic 4D World Modeling from LiDAR SequencesAlan Liang, Youquan Liu, Yu Yang, Dongyue Lu et al.AAAI 2026 · 12 citations
- VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric ControlSixiao Zheng, Minghao Yin, Wenbo Hu, Xiaoyu Li et al.CVPR 2026 · 27 citations
- TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion ModelsMark Yu, Wenbo Hu, Jinbo Xing, Ying ShanICCV 2025 · 25 citations
- LidaRF: Delving into Lidar for Neural Radiance Field on Street ScenesShanlin Sun, Bingbing Zhuang, Ziyu Jiang, Buyu Liu et al.CVPR 2024 · 11 citations
