Language Prompt for Autonomous Driving
Dongming Wu, Wencheng Han, Yingfei Liu, Tiancai Wang, Cheng-Zhong Xu, Xiangyu Zhang, Jianbing Shen
Abstract
A new trend in the computer vision community is to capture objects of interest following flexible human command represented by a natural language prompt. However, the progress of using language prompts in driving scenarios is stuck in a bottleneck due to the scarcity of paired prompt-instance data. To address this challenge, we propose the first objectcentric language prompt set for driving scenes within 3D, multi-view, and multi-frame space, named NuPrompt. It expands nuScenes dataset by constructing a total of 40,147 language descriptions, each referring to an average of 7.4 object tracklets. Based on the object-text pairs from the new benchmark, we formulate a novel prompt-based driving task, i.e., employing a language prompt to predict the described object trajectory across views and frames. Furthermore, we provide a simple end-to-end baseline model based on Transformer, named PromptTrack. Experiments show that our Prompt-Track achieves impressive performance on NuPrompt. We hope this work can provide some new insights for the selfdriving community. The data and code have been released at https://github.com/wudongming97/Prompt4Driving .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ef1b82d-90b9-4061-a932-7fa8c189ca93Cited by top-tier papers23
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningZewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang et al.NeurIPS 2025 · 310 citations
- SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic SegmentationBin Xie, Jiale Cao, Jin Xie, Fahad Shahbaz Khan et al.CVPR 2024 · 57 citations
- AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the WildXiaolou Sun, Wufei Si, Wenhui Ni, Yuntian Li et al.ICLR 2026 · 25 citations
- Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric PerspectivesShaoyuan Xie, Lingdong Kong, Yuhao Dong, Chonghao Sima et al.ICCV 2025 · 25 citations
- SpaceVista: All-Scale Visual Spatial Reasoning from mm to kmPeiwen Sun, Shiqiang Lang, Dongming Wu, Ding Yi et al.ICML 2026 · 19 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object DetectionShihao Wang, Yingfei Liu, Tiancai Wang, Ying Li et al.ICCV 2023 · 399 citations
- NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving ScenarioTianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao et al.AAAI 2024 · 314 citations
- OnlineRefer: A Simple Online Baseline for Referring Video Object SegmentationDongming Wu, Tiancai Wang, Yuang Zhang, Xiangyu Zhang et al.ICCV 2023 · 82 citations
- Type-to-Track: Retrieve Any Object via Prompt-based TrackingPha A. Nguyen, Kha Gia Quach, Kris Kitani, Khoa LuuNeurIPS 2023 · 38 citations
Related papers
- SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous DrivingPeizheng Li, Zhenghao Zhang, David Holtz, Hang Yu et al.CVPR 2026 · 32 citations
- Open-vocabulary object 6D pose estimationJaime Corsetti, Davide Boscaini, Changjae Oh, Andrea Cavallaro et al.CVPR 2024
- CXTrack: Improving 3D Point Cloud Tracking with Contextual InformationTian-Xing Xu, Yuan-Chen Guo, Yu-Kun Lai, Song-Hai ZhangCVPR 2023
- Towards More Flexible and Accurate Object Tracking With Natural Language: Algorithms and BenchmarkXiao Wang, Xiujun Shu, Zhipeng Zhang, Bo Jiang et al.CVPR 2021
- DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion AlignmentXiaofan Li, Chenming Wu, Zhao Yang, Zhihao Xu et al.ACM MM 2025
