MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
Ruiyuan Gao, Kai Chen, Bo Xiao, Lanqing Hong, Zhenguo Li, Qiang Xu
Abstract
The rapid advancement of diffusion models has greatly improved video synthesis, especially in controllable video generation, which is vital for applications like autonomous driving. Although DiT with 3D VAE has become a standard framework for video generation, it introduces challenges in controllable driving video generation, especially for framewise geometric control, rendering existing methods ineffective. To address these issues, we propose MagicDriveV2, a novel approach that integrates the MVDiT block and spatial-temporal conditional encoding to enable multiview video generation and precise geometric control. Additionally, we introduce an efficient method for obtaining contextual descriptions for videos to support diverse textual control, along with a progressive training strategy using mixed video data to enhance training efficiency and generalizability. Consequently, MagicDrive-V2 enables multi-view driving video synthesis with resolution and frame count (compared to current SOTA), rich contextual control, and geometric controls. Extensive experiments demonstrate MagicDrive-V2's ability, unlocking broader applications in autonomous driving. Project page: flymin.github.io/magicdrive-v2/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 339fd10a-9e47-4530-8378-ec507ab66fabCited by top-tier papers19
- DriveLaW: Unifying Planning and Video Generation in a Latent Driving WorldTianze Xia, Yongkang Li, Lijun Zhou, Jingfeng Yao et al.CVPR 2026 · 58 citations
- WorldLens: Full-Spectrum Evaluations of Driving World Models in Real WorldAo Liang, Lingdong Kong, Tianyi Yan, Hongsi Liu et al.CVPR 2026 · 28 citations
- Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal ConsistencyXiangyu Guo, Zhanqian Wu, Kaixin Xiong, Ziyang Xu et al.NeurIPS 2025 · 24 citations
- Rethinking Video Generation Model for the Embodied WorldYufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li et al.ICML 2026 · 24 citations
- X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible ControllabilityYu Yang, Alan Liang, Jianbiao Mei, Yukai Ma et al.NeurIPS 2025 · 22 citations
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- MagicDrive: Street View Generation with Diverse 3D Geometry ControlRuiyuan Gao, Kai Chen, Enze Xie, Lanqing Hong et al.ICLR 2024 · 248 citations
- MAD: Motion Appearance Decoupling for efficient Driving World ModelsAhmad Rahimi, Valentin Gerard, Eloi Zablocki, Matthieu Cord et al.CVPR 2026 · 7 citations
- UniMLVG: Unified Framework for Multi-View Long Video Generation with Comprehensive Control Capabilities for Autonomous DrivingRui Chen, Zehuan Wu, Yichen Liu, Yuxin Guo et al.ICCV 2025 · 2 citations
- GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video GenerationZhenya Yang, Zhe Liu, Yuxiang Lu, Liping Hou et al.CVPR 2026 · 12 citations
- ConsistentCity: Semantic Flow-Guided Occupancy DiT for Temporally Consistent Driving Scene SynthesisBenjin Zhu, Xiaogang Wang, Hongsheng LiICCV 2025 · 1 citation
