Bird's-eye-view Informed Reasoning Driver
Yinuo Wang, Mining Tan, Yuanxin Zhong, Wang zhitao, Siyuan Cheng
Abstract
Motion planning in complex environments remains a core challenge for autonomous driving. While existing rule-based or imitation learning-based motion planning methods perform well in common scenarios, they often struggle with complex, long-tail scenarios. To address this problem, we introduce the Bird's-eye-view Informed Reasoning Driver (BIRDriver), a hierarchical framework that combines a Vision-Language Model (VLM) with a motion planner. BIRDriver leverages the commonsense reasoning capabilities of the VLM to effectively handle these challenging long-tail scenarios. Unlike prior methods that require domain-specific encoders and costly alignment, our approach compresses the environment into a single-frame bird's-eye-view (BEV) map, a paradigm that enables the model to fully leverage its knowledge from internet-scale pre-training. It then generates high-level key points, which are encoded and passed to the motion planner to produce the final trajectory. However, a major challenge is that standard VLMs struggle to generate the precise numerical coordinates required for such key points. We address this limitation by fine-tuning them on a composite dataset of three auxiliary types to enhance spatial localization, scene understanding, and key-point generation, complemented by a token-level weighted mechanism for improved numerical precision. Experiments on the nuPlan dataset demonstrate that BIRDriver outperforms the base motion planner in most cases on both Test14-hard and Test14-random benchmarks, and achieves state-of-theart (SOTA) performance on the InterPlan long-tail benchmark. Weighted SFT Loss Traffic Light Road Info Agents System prompts: Defining the role of the driver; The principles of safety, efficiency, and comfort; Elements and meanings of BEV map. VLM Input Bird's-Eye-View Map User prompts: Last moment ego speed; Commands for generating key points SFT dataset Spatial Localization Driving Scene CoT KeyPoints
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04e1363e-e622-43df-9d39-a07602c683fbBuilds on11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao et al.ICCV 2023 · 602 citations
- LoRA+: Efficient Low Rank Adaptation of Large ModelsSoufiane Hayou, Nikhil Ghosh, Bin YuICML 2024 · 388 citations
- GameFormer: Game-theoretic Modeling and Learning of Transformer-based Interactive Prediction and Planning for Autonomous DrivingZhiyu Huang, Haochen Liu, Chen LvICCV 2023 · 209 citations
Related papers
- VLMPlanner: Integrating Visual Language Models with Motion PlanningZhipeng Tang, Sha Zhang, Jiajun Deng, Chenjie Wang et al.ACM MM 2025 · 2 citations
- DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous DrivingLingjun Zhang, Changjie Wu, Linzhe Shi, Jiangyang Li et al.ICML 2026
- VLP: Vision Language Planning for Autonomous DrivingChenbin Pan, Burhaneddin Yaman, Tommaso Nesti, Abhirup Mallik et al.CVPR 2024 · 53 citations
- SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Drivingjingyu li, Junjie Wu, Dongnan Hu, Xiangkai Huang et al.CVPR 2026 · 36 citations
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous DrivingJianhua Han, Meng Tian, Jiangtong Zhu, Fan He et al.CVPR 2026 · 10 citations
