Lune

NeurIPS2025Top-tier venue

Overcoming Challenges of Long-Horizon Prediction in Driving World Models

Arian Mousakhan, Sudhanshu Mittal, Silvio Galesso, Karim Farid, Thomas Brox

2025Year
1Citations

Abstract

Existing world models for autonomous driving struggle with long-horizon generation and generalization to challenging scenarios. In this work, we develop a model using simple design choices, and without additional supervision or sensors, such as maps, depth, or multiple cameras. We show that our model yields state-of-the-art performance, despite having only 469M parameters and being trained on 280h of video data. It particularly stands out in difficult scenarios like turning maneuvers and urban traffic. We test whether discrete token models possibly have advantages over continuous models based on flow matching. To this end, we set up a hybrid tokenizer that is compatible with both approaches and allows for a side-by-side comparison. Our study concludes in favor of the continuous autoregressive model, which is less brittle on individual design choices and more powerful than the model built on discrete tokens. Project page with code, model checkpoints and visualization can be found here: https://lmb-freiburg.github.io/orbis.github.io

2 Related Work World Models. The ability of world models to do real-world simulation can be useful for policy learning [59,26], sample efficient RL [23,24,79,42], and representation learning [89]. Previous world models have been limited to gaming [24,23,21] and other simulated environments [12]. Recent breakthroughs in video generative modeling [83,6] have led to future video prediction models -an essential building block for world models.

Multiple driving world models [75,77,44] use BEV (Bird's-Eye-View) annotations like depth maps, 3D bounding boxes, road maps to generate new scenarios. DriveDreamer [75] incorporated multimodal input, such as traffic conditions, text prompts, and driving actions, for future frames and action generation. Many other works [77,90,76] extended this idea to multi-view video generation. Some recent works [90,32] also use LLMs and VLMs [47] for spatial reasoning. Although these models show high quality generation, they rely on heavy external knowledge. Such heavy reliance limits the model's ability to generalize to new environments. In this work, we train a generalizable world model using unannotated front-camera videos and only fine-tune for ego-motion control.

Recent driving world models [36,82,17,1,25,30,66] trained predominantly on raw driving video data have shown the ability to simulate realistic future scenes in unseen environments. DriveGAN [36], among the first works to train on real-world driving data, showed realistic future generation with ego-motion and environment controllability. GAIA-1 [30] further enhanced the quality of future prediction and added controllability through text, in addition to action input. Diffusion-based world models [17,25] fine-tuned general-purpose pre-trained video generation models like SVD [6] to produce future video predictions at high resolution and high frame rate. Driving world models -Vista [17] and GEM [25] demonstrate high-quality rollouts up to 15 seconds. DrivingWorld [31] further enables longer and more coherent rollouts.

Generative models based on vector-quantized tokens like autoregressive [81,38] and masked generative models [87,48,20], have also demonstrated strong performance in video generation due to their strong capability in modeling dynamics and representation learning. For world modeling, Genie [7] and GAIA-1 [30] have demonstrated generalized world modeling capabilities with interactive control and long-horizon rollouts respectively. We also observe that quantized driving world model can perform long-horizon generation.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ffee0d5a-f897-4361-8874-9be9c40cfe77

Builds on45

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines