NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario
Tianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao, Yu-Gang Jiang
摘要
We introduce a novel visual question answering (VQA) task in the context of autonomous driving, aiming to answer natural language questions based on street-view clues. Compared to traditional VQA tasks, VQA in autonomous driving scenario presents more challenges. Firstly, the raw visual data are multi-modal, including images and point clouds captured by camera and LiDAR, respectively. Secondly, the data are multi-frame due to the continuous, real-time acquisition. Thirdly, the outdoor scenes exhibit both moving foreground and static background. Existing VQA benchmarks fail to adequately address these complexities. To bridge this gap, we propose NuScenes-QA, the first benchmark for VQA in the autonomous driving scenario, encompassing 34K visual scenes and 460K question-answer pairs. Specifically, we leverage existing 3D detection annotations to generate scene graphs and design question templates manually. Subsequently, the question-answer pairs are generated programmatically based on these templates. Comprehensive statistics prove that our NuScenes-QA is a balanced large-scale benchmark with diverse question formats. Built upon it, we develop a series of baselines that employ advanced 3D detection and VQA techniques. Our extensive experiments highlight the challenges posed by this new task. Codes and dataset are available at https://github.com/qiantianwen/NuScenes-QA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper74
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningZewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang 等NeurIPS 2025 · 被引用 310 次
- ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous DrivingYongkang Li, Kaixin Xiong, Xiangyu Guo, Fang Li 等ICLR 2026 · 被引用 196 次
- Language Prompt for Autonomous DrivingDongming Wu, Wencheng Han, Yingfei Liu, Tiancai Wang 等AAAI 2025 · 被引用 150 次
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelXingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma 等AAAI 2026 · 被引用 119 次
- SimGen: Simulator-conditioned Driving Scene GenerationYunsong Zhou, Michael Simon, Zhenghao Mark Peng, Sicheng Mo 等NeurIPS 2024 · 被引用 44 次
它引用的顶会 Paper8
- ScanQA: 3D Question Answering for Spatial Scene UnderstandingDaichi Azuma, Taiki Miyanishi, Shuhei Kurita, Motoaki KawanabeCVPR 2022 · 被引用 135 次
- Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question AnsweringJianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao 等AAAI 2020 · 被引用 129 次
- Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting LearningYang Jiao, Zequn Jie, Shaoxiang Chen, Lechao Cheng 等AAAI 2024 · 被引用 13 次
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora 等CVPR 2020
- Center-Based 3D Object Detection and TrackingTianwei Yin, Xingyi Zhou, Philipp KrähenbühlCVPR 2021
相关 Paper
- NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language ModelsSung-Yeon Park, Can Cui, Yunsheng Ma, Ahmadreza Moradipari 等ICCV 2025 · 被引用 6 次
- OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object DetectionAdrian Chow, Evelien Riddell, Yimu Wang, Sean Sedwards 等ICCV 2025 · 被引用 2 次
- Embodied Scene Understanding for Vision Language Models via MetaVQAWeizhen Wang, Chenda Duan, Zhenghao Peng, Yuxin Liu 等CVPR 2025
- PointAugmenting: Cross-Modal Augmentation for 3D Object DetectionChunwei Wang, Chao Ma, Ming Zhu, Xiaokang YangCVPR 2021
- 3D Question Answering for City Scene UnderstandingPenglei Sun, Yaoxian Song, Xiang Liu, Xiaofei Yang 等ACM MM 2024 · 被引用 6 次
