Fine-flow Distilling Coarse-flow Video Generation for Long-Term Driving World Model
Xiaodong Wang, Zhirong Wu, Peixi Peng
Abstract
Driving world models are used to simulate futures by video generation based on the condition of the current state and actions. However, current models often suffer serious error accumulations when predicting the long-term future, which limits practical applications. Recent studies utilize the Diffusion Transformer (DiT) as the backbone of driving world models to improve learning flexibility. However, these models are always trained on short video clips, and multiple roll-out generations struggle to produce consistent and reasonable long videos due to the training-inference gap. To this end, we propose several solutions to build a simple yet effective long-term driving world model. First, we hierarchically decouple world model learning into large motion learning and bidirectional continuous motion learning. Then, considering the continuity of driving scenes, we propose a simple distillation method where fine-grained video flows are self-supervised signals for coarse-grained flows. The distillation is designed to improve the coherence of infinite video generation. The coarse-grained and fine-grained modules are coordinated to generate long-term and temporally coherent videos. On NuScenes, compared with the state-of-the-art front-view models, our model improves FVD by 27% and reduces inference time by 85% for the video task of generating 110+ frames.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2861a0a8-23b2-440a-8aa3-c12f50cb0c12Builds on26
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
Related papers
- DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion AlignmentXiaofan Li, Chenming Wu, Zhao Yang, Zhihao Xu et al.ACM MM 2025
- MaskGWM: A Generalizable Driving World Model with Video Mask ReconstructionJingcheng Ni, Yuxin Guo, Yichen Liu, Rui Chen et al.CVPR 2025
- DriveLaW: Unifying Planning and Video Generation in a Latent Driving WorldTianze Xia, Yongkang Li, Lijun Zhou, Jingfeng Yao et al.CVPR 2026 · 58 citations
- Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent SpaceJian Zhu, Zhengyu Jia, Tian Gao, Jiaxin Deng et al.AAAI 2026 · 5 citations
- Astra: General Interactive World Model with Autoregressive DenoisingYixuan Zhu, Jiaqi Feng, Wenzhao Zheng, Yuan Gao et al.ICLR 2026 · 29 citations
