CorrectionPlanner: Self-Correction Planner with Reinforcement Learning in Autonomous Driving
Yihong Guo, Dongqiangzi Ye, Sijia Chen, Anqi Liu, Xianming Liu
摘要
Autonomous driving requires safe planning, but most learning-based planners lack explicit selfcorrection ability: once an unsafe action is proposed, there is no mechanism to correct it. Thus, we propose CorrectionPlanner, an autoregressive planner with self-correction that models planning as motion-token generation within a propose, evaluate, and correct loop. At each planning step, the policy proposes an action, namely a motion token, and a learned collision critic predicts whether it will induce a collision within a short horizon. If the critic predicts a collision, we retain the sequence of historical unsafe motion tokens as a self-correction trace, generate the next motion token conditioned on it, and repeat this process until the safe motion token is proposed or the safety criterion is met. This self-correction trace, consisting of all the unsafe motion tokens, represents the planner's correction process in motion-token space (analogous to reasoning trace in language models). We train the planner with imitation learning followed by model-based reinforcement learning using rollouts from a pretrained world model that realistically models agents' reactive behaviors. Closed-loop evaluations show that Correc-tionPlanner reduces the collision rate by over 20% on Waymax and obtains state-of-the-art planning scores on nuPlan.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion DatasetScott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu 等ICCV 2021 · 被引用 817 次
- GameFormer: Game-theoretic Modeling and Learning of Transformer-based Interactive Prediction and Planning for Autonomous DrivingZhiyu Huang, Haochen Liu, Chen LvICCV 2023 · 被引用 209 次
- Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMsXuan Zhang, Chao Du, Tianyu Pang, Qian Liu 等NeurIPS 2024 · 被引用 177 次
- Sample Efficient Reinforcement Learning with REINFORCEJunzi Zhang, Jongho Kim, Brendan O'Donoghue, Stephen P. BoydAAAI 2021 · 被引用 162 次
- Forecast-MAE: Self-supervised Pre-training for Motion Forecasting with Masked AutoencodersJie Cheng, Xiaodong Mei, Ming LiuICCV 2023 · 被引用 123 次
相关 Paper
- CarPlanner: Consistent Auto-regressive Trajectory Planning for Large-Scale Reinforcement Learning in Autonomous DrivingDongkun Zhang, Jiaming Liang, Ke Guo, Sha Lu 等CVPR 2025
- Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive ReasoningZhenghao Peng, Wenhao Ding, Yurong You, Yuxiao Chen 等CVPR 2026 · 被引用 25 次
- Flow Matching-Based Autonomous Driving Planning with Advanced Interactive Behavior ModelingTianyi Tan, Yinan Zheng, Ruiming Liang, Zexu Wang 等NeurIPS 2025 · 被引用 36 次
- Diffusion-Based Planning for Autonomous Driving with Flexible GuidanceYinan Zheng, Ruiming Liang, Kexin Zheng, Jinliang Zheng 等ICLR 2025
- DrivingGPT: Unifying Driving World Modeling and Planning with Multi-Modal Autoregressive TransformersYuntao Chen, Yuqi Wang, Zhaoxiang ZhangICCV 2025 · 被引用 7 次
