MAD: Motion Appearance Decoupling for efficient Driving World Models
Ahmad Rahimi, Valentin Gerard, Eloi Zablocki, Matthieu Cord, Alexandre Alahi
摘要
Recent video diffusion models generate photorealistic, temporally coherent videos, yet they fall short as reliable world models for autonomous driving, where structured motion and physically consistent interactions are essential. Adapting these generalist video models to driving domains has shown promise but typically requires massive domain-specific data and costly fine-tuning. We propose an efficient adaptation framework that converts generalist video diffusion models into controllable driving world models with minimal supervision. The key idea is to decouple motion learning from appearance synthesis. First, the model is adapted to predict structured motion in a simplified form: videos of skeletonized agents and scene elements, focusing learning on physical and social plausibility. Then, the same backbone is reused to synthesize realistic RGB videos conditioned on these motion sequences, effectively"dressing"the motion with texture and lighting. This two-stage process mirrors a reasoning-rendering paradigm: first infer dynamics, then render appearance. Our experiments show this decoupled approach is exceptionally efficient: adapting SVD, we match prior SOTA models with less than 6% of their compute. Scaling to LTX, our MAD-LTX model outperforms all open-source competitors, and supports a comprehensive suite of text, ego, and object controls. Project page: https://vita-epfl.github.io/MAD-World-Model/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta 等NeurIPS 2024 · 被引用 403 次
- Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion ModelingXiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian 等SIGGRAPH 2024 · 被引用 66 次
- ReSim: Reliable World Simulation for Autonomous DrivingJiazhi Yang, Kashyap Chitta, Shenyuan Gao, Long Chen 等NeurIPS 2025 · 被引用 53 次
- Generalized Predictive Model for Autonomous DrivingJiazhi Yang, Shenyuan Gao, Yihang Qiu, Li Chen 等CVPR 2024 · 被引用 32 次
相关 Paper
- Vid2World: Crafting Video Diffusion Models to Interactive World ModelsSiqiao Huang, Jialong Wu, Qixing Zhou, Shangchen Miao 等ICLR 2026 · 被引用 68 次
- Epona: Autoregressive Diffusion World Model for Autonomous DrivingKaiwen Zhang, Zhenyu Tang, Xiaotao Hu, Xingang Pan 等ICCV 2025 · 被引用 14 次
- World-consistent Video Diffusion with Explicit 3D ModelingQihang Zhang, Shuangfei Zhai, Miguel Ángel Bautista Martin, Kevin Miao 等CVPR 2025
- LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video DiffusionYisu Zhang, Chenjie Cao, Chaohui Yu, Jianke ZhuICCV 2025 · 被引用 2 次
- MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive ControlRuiyuan Gao, Kai Chen, Bo Xiao, Lanqing Hong 等ICCV 2025 · 被引用 11 次
