AdaWM: Adaptive World Model based Planning for Autonomous Driving
Hang Wang, Xin Ye, Feng Tao, Chenbin Pan, Abhirup Mallik, Burhaneddin Yaman, Liu Ren, Junshan Zhang
摘要
World model based reinforcement learning (RL) has emerged as a promising approach for autonomous driving, which learns a latent dynamics model and uses it to train a planning policy. To speed up the learning process, the pretrain-finetune paradigm is often used, where online RL is initialized by a pretrained model and a policy learned offline. However, naively performing such initialization in RL may result in dramatic performance degradation during the online interactions in the new task. To tackle this challenge, we first analyze the performance degradation and identify two primary root causes therein: the mismatch of the planning policy and the mismatch of the dynamics model, due to distribution shift. We further analyze the effects of these factors on performance degradation during finetuning, and our findings reveal that the choice of finetuning strategies plays a pivotal role in mitigating these effects. We then introduce AdaWM, an Adaptive World Model based planning method, featuring two key steps: (a) mismatch identification, which quantifies the mismatches and informs the finetuning strategy, and (b) alignment-driven finetuning, which selectively updates either the policy or the model as needed using efficient low-rank updates. Extensive experiments on the challenging CARLA driving tasks demonstrate that AdaWM significantly improves the finetuning process, resulting in more robust and efficient performance in autonomous driving systems. 1 INTRODUCTION Automated vehicles (AVs) are poised to revolutionize future mobility systems with enhanced safety and efficiency Yurtsever et al. (2020); Kalra & Paddock (2016); Maurer et al. (2016). Despite significant progress Teng et al. (2023); Hu et al. (2023); Jiang et al. (2023), developing AVs capable of navigating complex, diverse real-world scenarios remains challenging, particularly in unforeseen situations Campbell et al. (2010); Chen et al. (2024). Autonomous vehicles must learn the complex dynamics of environments, predict future scenarios accurately and swiftly, and take timely actions such as emergency braking. Thus motivated, in this work, we devise adaptive world model to advance embodied AI and improve the planning capability of autonomous driving systems. World model (WM) based reinforcement learning (RL) has emerged as a promising self-supervised approach for autonomous driving Chen et al. (2024); Wang et al. (2024); Guan et al. (2024); Li et al. (2024). This end-to-end method maps sensory inputs directly to control outputs, offering improved efficiency and robustness over traditional modular architectures Yurtsever et al. (2020); Chen et al. (2024). By learning a latent dynamics model from observations and actions, the system can predict future events and optimize policy decisions, enhancing generalization across diverse environments. Recent models like DreamerV2 and DreamerV3 have demonstrated strong performance across both 2D and 3D environments Hafner et al. (2020; 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail ScenariosRunsheng Xu, Hubert Lin, Wonseok Jeon, Hao Feng 等CVPR 2026 · 被引用 82 次
- Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)Zhenjie Yang, Xiaosong Jia, Qifeng Li, Xue Yang 等NeurIPS 2025 · 被引用 65 次
- Argus: Resilience-Oriented Safety Assurance Framework for End-to-End ADSsDingji Wang, You Lu, Bihuan Chen, Shuo Hao 等ASE 2025
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao 等ICCV 2023 · 被引用 602 次
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
- Jump-Start Reinforcement LearningIkechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu 等ICML 2023 · 被引用 158 次
相关 Paper
- AdaWorld: Learning Adaptable World Models with Latent ActionsShenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang 等ICML 2025
- WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous DrivingPengxuan Yang, Ben Lu, Zhongpu Xia, Chao Han 等AAAI 2026 · 被引用 8 次
- Iso-Dream: Isolating and Leveraging Noncontrollable Visual Dynamics in World ModelsMinting Pan, Xiangming Zhu, Yunbo Wang, Xiaokang YangNeurIPS 2022 · 被引用 74 次
- WPT: World-to-Policy Transfer via Online World Model DistillationGuangfeng Jiang, Yueru Luo, Jun Liu, Yi Huang 等CVPR 2026 · 被引用 4 次
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningZewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang 等NeurIPS 2025 · 被引用 310 次
