Overcoming Challenges of Long-Horizon Prediction in Driving World Models
Arian Mousakhan, Sudhanshu Mittal, Silvio Galesso, Karim Farid, Thomas Brox
摘要
Existing world models for autonomous driving struggle with long-horizon generation and generalization to challenging scenarios. In this work, we develop a model using simple design choices, and without additional supervision or sensors, such as maps, depth, or multiple cameras. We show that our model yields state-of-the-art performance, despite having only 469M parameters and being trained on 280h of video data. It particularly stands out in difficult scenarios like turning maneuvers and urban traffic. We test whether discrete token models possibly have advantages over continuous models based on flow matching. To this end, we set up a hybrid tokenizer that is compatible with both approaches and allows for a side-by-side comparison. Our study concludes in favor of the continuous autoregressive model, which is less brittle on individual design choices and more powerful than the model built on discrete tokens. Project page with code, model checkpoints and visualization can be found here: https://lmb-freiburg.github.io/orbis.github.io
2 Related Work World Models. The ability of world models to do real-world simulation can be useful for policy learning [59,26], sample efficient RL [23,24,79,42], and representation learning [89]. Previous world models have been limited to gaming [24,23,21] and other simulated environments [12]. Recent breakthroughs in video generative modeling [83,6] have led to future video prediction models -an essential building block for world models.
Multiple driving world models [75,77,44] use BEV (Bird's-Eye-View) annotations like depth maps, 3D bounding boxes, road maps to generate new scenarios. DriveDreamer [75] incorporated multimodal input, such as traffic conditions, text prompts, and driving actions, for future frames and action generation. Many other works [77,90,76] extended this idea to multi-view video generation. Some recent works [90,32] also use LLMs and VLMs [47] for spatial reasoning. Although these models show high quality generation, they rely on heavy external knowledge. Such heavy reliance limits the model's ability to generalize to new environments. In this work, we train a generalizable world model using unannotated front-camera videos and only fine-tune for ego-motion control.
Recent driving world models [36,82,17,1,25,30,66] trained predominantly on raw driving video data have shown the ability to simulate realistic future scenes in unseen environments. DriveGAN [36], among the first works to train on real-world driving data, showed realistic future generation with ego-motion and environment controllability. GAIA-1 [30] further enhanced the quality of future prediction and added controllability through text, in addition to action input. Diffusion-based world models [17,25] fine-tuned general-purpose pre-trained video generation models like SVD [6] to produce future video predictions at high resolution and high frame rate. Driving world models -Vista [17] and GEM [25] demonstrate high-quality rollouts up to 15 seconds. DrivingWorld [31] further enables longer and more coherent rollouts.
Generative models based on vector-quantized tokens like autoregressive [81,38] and masked generative models [87,48,20], have also demonstrated strong performance in video generation due to their strong capability in modeling dynamics and representation learning. For world modeling, Genie [7] and GAIA-1 [30] have demonstrated generalized world modeling capabilities with interactive control and long-horizon rollouts respectively. We also observe that quantized driving world model can perform long-horizon generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper45
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Epona: Autoregressive Diffusion World Model for Autonomous DrivingKaiwen Zhang, Zhenyu Tang, Xiaotao Hu, Xingang Pan 等ICCV 2025 · 被引用 14 次
- Fine-flow Distilling Coarse-flow Video Generation for Long-Term Driving World ModelXiaodong Wang, Zhirong Wu, Peixi PengAAAI 2026
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta 等NeurIPS 2024 · 被引用 403 次
- DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video GenerationGuosheng Zhao, Xiaofeng Wang, Zheng Zhu, Xinze Chen 等AAAI 2025 · 被引用 31 次
- From Forecasting to Planning: Policy World Model for Collaborative State-Action PredictionZhida Zhao, Talas Fu, Yifan Wang, Lijun Wang 等NeurIPS 2025 · 被引用 36 次
