Lune

ICLR2026Top-tier venue

Bird's-eye-view Informed Reasoning Driver

Yinuo Wang, Mining Tan, Yuanxin Zhong, Wang zhitao, Siyuan Cheng

2026Year

Abstract

Motion planning in complex environments remains a core challenge for autonomous driving. While existing rule-based or imitation learning-based motion planning methods perform well in common scenarios, they often struggle with complex, long-tail scenarios. To address this problem, we introduce the Bird's-eye-view Informed Reasoning Driver (BIRDriver), a hierarchical framework that combines a Vision-Language Model (VLM) with a motion planner. BIRDriver leverages the commonsense reasoning capabilities of the VLM to effectively handle these challenging long-tail scenarios. Unlike prior methods that require domain-specific encoders and costly alignment, our approach compresses the environment into a single-frame bird's-eye-view (BEV) map, a paradigm that enables the model to fully leverage its knowledge from internet-scale pre-training. It then generates high-level key points, which are encoded and passed to the motion planner to produce the final trajectory. However, a major challenge is that standard VLMs struggle to generate the precise numerical coordinates required for such key points. We address this limitation by fine-tuning them on a composite dataset of three auxiliary types to enhance spatial localization, scene understanding, and key-point generation, complemented by a token-level weighted mechanism for improved numerical precision. Experiments on the nuPlan dataset demonstrate that BIRDriver outperforms the base motion planner in most cases on both Test14-hard and Test14-random benchmarks, and achieves state-of-theart (SOTA) performance on the InterPlan long-tail benchmark. Weighted SFT Loss Traffic Light Road Info Agents System prompts: Defining the role of the driver; The principles of safety, efficiency, and comfort; Elements and meanings of BEV map. VLM Input Bird's-Eye-View Map User prompts: Last moment ego speed; Commands for generating key points SFT dataset Spatial Localization Driving Scene CoT KeyPoints

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 04e1363e-e622-43df-9d39-a07602c683fb

Builds on11

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines