Bird's-eye-view Informed Reasoning Driver
Yinuo Wang, Mining Tan, Yuanxin Zhong, Wang zhitao, Siyuan Cheng
摘要
Motion planning in complex environments remains a core challenge for autonomous driving. While existing rule-based or imitation learning-based motion planning methods perform well in common scenarios, they often struggle with complex, long-tail scenarios. To address this problem, we introduce the Bird's-eye-view Informed Reasoning Driver (BIRDriver), a hierarchical framework that combines a Vision-Language Model (VLM) with a motion planner. BIRDriver leverages the commonsense reasoning capabilities of the VLM to effectively handle these challenging long-tail scenarios. Unlike prior methods that require domain-specific encoders and costly alignment, our approach compresses the environment into a single-frame bird's-eye-view (BEV) map, a paradigm that enables the model to fully leverage its knowledge from internet-scale pre-training. It then generates high-level key points, which are encoded and passed to the motion planner to produce the final trajectory. However, a major challenge is that standard VLMs struggle to generate the precise numerical coordinates required for such key points. We address this limitation by fine-tuning them on a composite dataset of three auxiliary types to enhance spatial localization, scene understanding, and key-point generation, complemented by a token-level weighted mechanism for improved numerical precision. Experiments on the nuPlan dataset demonstrate that BIRDriver outperforms the base motion planner in most cases on both Test14-hard and Test14-random benchmarks, and achieves state-of-theart (SOTA) performance on the InterPlan long-tail benchmark. Weighted SFT Loss Traffic Light Road Info Agents System prompts: Defining the role of the driver; The principles of safety, efficiency, and comfort; Elements and meanings of BEV map. VLM Input Bird's-Eye-View Map User prompts: Last moment ego speed; Commands for generating key points SFT dataset Spatial Localization Driving Scene CoT KeyPoints
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao 等ICCV 2023 · 被引用 602 次
- LoRA+: Efficient Low Rank Adaptation of Large ModelsSoufiane Hayou, Nikhil Ghosh, Bin YuICML 2024 · 被引用 388 次
- GameFormer: Game-theoretic Modeling and Learning of Transformer-based Interactive Prediction and Planning for Autonomous DrivingZhiyu Huang, Haochen Liu, Chen LvICCV 2023 · 被引用 209 次
相关 Paper
- VLMPlanner: Integrating Visual Language Models with Motion PlanningZhipeng Tang, Sha Zhang, Jiajun Deng, Chenjie Wang 等ACM MM 2025 · 被引用 2 次
- DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous DrivingLingjun Zhang, Changjie Wu, Linzhe Shi, Jiangyang Li 等ICML 2026
- VLP: Vision Language Planning for Autonomous DrivingChenbin Pan, Burhaneddin Yaman, Tommaso Nesti, Abhirup Mallik 等CVPR 2024 · 被引用 53 次
- SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Drivingjingyu li, Junjie Wu, Dongnan Hu, Xiangkai Huang 等CVPR 2026 · 被引用 36 次
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous DrivingJianhua Han, Meng Tian, Jiangtong Zhu, Fan He 等CVPR 2026 · 被引用 10 次
