DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
Shuyao Shang, Yuntao Chen, Yuqi Wang, Yingyan Li, Zhao-Xiang Zhang
摘要
End-to-end autonomous driving has substantially progressed by directly predicting future trajectories from raw perception inputs, which bypasses traditional modular pipelines. However, mainstream methods trained via imitation learning suffer from critical safety limitations, as they fail to distinguish between trajectories that appear human-like but are potentially unsafe. Some recent approaches attempt to address this by regressing multiple rule-driven scores but decoupling supervision from policy optimization, resulting in suboptimal performance. To tackle these challenges, we propose DriveDPO, a Safety Direct Preference Optimization Policy Learning framework. First, we distill a unified policy distribution from human imitation similarity and rule-based safety scores for direct policy optimization. Further, we introduce an iterative Direct Preference Optimization stage formulated as trajectory-level preference alignment. Extensive experiments on the NAVSIM benchmark demonstrate that DriveDPO achieves a new state-of-the-art PDMS of 90.0. Furthermore, qualitative results across diverse challenging scenarios highlight DriveDPO's ability to produce safer and more reliable driving behaviors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous DrivingPeizheng Li, Zhenghao Zhang, David Holtz, Hang Yu 等CVPR 2026 · 被引用 32 次
- SafeDrive: Fine-Grained Safety Reasoning for End-to-End Driving in a Sparse WorldJungho Kim, Jiyong Oh, Seunghoon Yu, Hongjae Shin 等CVPR 2026 · 被引用 8 次
- WPT: World-to-Policy Transfer via Online World Model DistillationGuangfeng Jiang, Yueru Luo, Jun Liu, Yi Huang 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper26
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky 等ICML 2024 · 被引用 973 次
相关 Paper
- DriveSuprim: Towards Precise Trajectory Selection for End-to-End PlanningWenhao Yao, Zhenxin Li, Shiyi Lan, Zi Wang 等AAAI 2026 · 被引用 46 次
- Prioritizing Perception-Guided Self-Supervision: A New Paradigm for Causal Modeling in End-to-End Autonomous DrivingYi Huang, Zhan Qu, Lihui Jiang, Bingbing Liu 等NeurIPS 2025 · 被引用 5 次
- DriveAdapter: Breaking the Coupling Barrier of Perception and Planning in End-to-End Autonomous DrivingXiaosong Jia, Yulu Gao, Li Chen, Junchi Yan 等ICCV 2023 · 被引用 154 次
- CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous DrivingXiaoji Zheng, Ziyuan Yang, Yanhao Chen, Yuhang PENG 等ICML 2026
- Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous DrivingPengxiang Li, Yinan Zheng, Yue Wang, Huimin Wang 等ICLR 2026 · 被引用 24 次
