Overcoming Challenges of Long-Horizon Prediction in Driving World Models
Arian Mousakhan, Sudhanshu Mittal, Silvio Galesso, Karim Farid, Thomas Brox
Abstract
Existing world models for autonomous driving struggle with long-horizon generation and generalization to challenging scenarios. In this work, we develop a model using simple design choices, and without additional supervision or sensors, such as maps, depth, or multiple cameras. We show that our model yields state-of-the-art performance, despite having only 469M parameters and being trained on 280h of video data. It particularly stands out in difficult scenarios like turning maneuvers and urban traffic. We test whether discrete token models possibly have advantages over continuous models based on flow matching. To this end, we set up a hybrid tokenizer that is compatible with both approaches and allows for a side-by-side comparison. Our study concludes in favor of the continuous autoregressive model, which is less brittle on individual design choices and more powerful than the model built on discrete tokens. Project page with code, model checkpoints and visualization can be found here: https://lmb-freiburg.github.io/orbis.github.io
2 Related Work World Models. The ability of world models to do real-world simulation can be useful for policy learning [59,26], sample efficient RL [23,24,79,42], and representation learning [89]. Previous world models have been limited to gaming [24,23,21] and other simulated environments [12]. Recent breakthroughs in video generative modeling [83,6] have led to future video prediction models -an essential building block for world models.
Multiple driving world models [75,77,44] use BEV (Bird's-Eye-View) annotations like depth maps, 3D bounding boxes, road maps to generate new scenarios. DriveDreamer [75] incorporated multimodal input, such as traffic conditions, text prompts, and driving actions, for future frames and action generation. Many other works [77,90,76] extended this idea to multi-view video generation. Some recent works [90,32] also use LLMs and VLMs [47] for spatial reasoning. Although these models show high quality generation, they rely on heavy external knowledge. Such heavy reliance limits the model's ability to generalize to new environments. In this work, we train a generalizable world model using unannotated front-camera videos and only fine-tune for ego-motion control.
Recent driving world models [36,82,17,1,25,30,66] trained predominantly on raw driving video data have shown the ability to simulate realistic future scenes in unseen environments. DriveGAN [36], among the first works to train on real-world driving data, showed realistic future generation with ego-motion and environment controllability. GAIA-1 [30] further enhanced the quality of future prediction and added controllability through text, in addition to action input. Diffusion-based world models [17,25] fine-tuned general-purpose pre-trained video generation models like SVD [6] to produce future video predictions at high resolution and high frame rate. Driving world models -Vista [17] and GEM [25] demonstrate high-quality rollouts up to 15 seconds. DrivingWorld [31] further enables longer and more coherent rollouts.
Generative models based on vector-quantized tokens like autoregressive [81,38] and masked generative models [87,48,20], have also demonstrated strong performance in video generation due to their strong capability in modeling dynamics and representation learning. For world modeling, Genie [7] and GAIA-1 [30] have demonstrated generalized world modeling capabilities with interactive control and long-horizon rollouts respectively. We also observe that quantized driving world model can perform long-horizon generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ffee0d5a-f897-4361-8874-9be9c40cfe77Builds on45
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Epona: Autoregressive Diffusion World Model for Autonomous DrivingKaiwen Zhang, Zhenyu Tang, Xiaotao Hu, Xingang Pan et al.ICCV 2025 · 14 citations
- Fine-flow Distilling Coarse-flow Video Generation for Long-Term Driving World ModelXiaodong Wang, Zhirong Wu, Peixi PengAAAI 2026
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta et al.NeurIPS 2024 · 403 citations
- DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video GenerationGuosheng Zhao, Xiaofeng Wang, Zheng Zhu, Xinze Chen et al.AAAI 2025 · 31 citations
- From Forecasting to Planning: Policy World Model for Collaborative State-Action PredictionZhida Zhao, Talas Fu, Yifan Wang, Lijun Wang et al.NeurIPS 2025 · 36 citations
