RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
Zhijian Huang, Chengjian Feng, Feng Yan, Baihui Xiao, Zequn Jie, Yujie Zhong, Xiaodan Liang, Lin Ma
摘要
Large Multimodal Models (LMMs) have demonstrated exceptional comprehension and interpretation capabilities in Autonomous Driving (AD) by incorporating large language models. Despite the advancements, current datadriven approaches tend to concentrate on a single dataset and specific tasks, neglecting their overall capabilities and ability to generalize. To bridge these gaps, we propose RoboTron-Drive, a general large multimodal model designed to process diverse data inputs, such as images and multi-view videos, while performing a broad spectrum of AD tasks, including perception, prediction, and planning. Initially, the model undergoes curriculum pretraining to process varied visual signals and perform basic visual comprehension and perception tasks. Subsequently, we augment and standardize various datasets to finetune the model, resulting in an all-in-one LMM for autonomous driving. To assess the general capabilities and generalization ability, we conduct evaluations on six public benchmarks and undertake zero-shot transfer on three unseen datasets, where RoboTron-Drive achieves state-of-theart performance across all tasks. We hope RoboTron-Drive as a promising solution for in the real world.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Drivingjingyu li, Junjie Wu, Dongnan Hu, Xiangkai Huang 等CVPR 2026 · 被引用 36 次
- VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous DrivingJie Wang, Guang Li, Zhijian Huang, Chenxu Dang 等CVPR 2026 · 被引用 20 次
- CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action ModelDapeng Zhang, Fei Shen, Rui Zhao, Yinda Chen 等NeurIPS 2025 · 被引用 8 次
- AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture -of-Transformers for End-to-End Autonomous DrivingWenhui (Oscar) Huang, Songyan Zhang, Qihang Huang, Zhidong Wang 等ICML 2026 · 被引用 6 次
- HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and GenerationXin Zhou, Dingkang Liang, Sifan Tu, Xiwu Chen 等ICCV 2025 · 被引用 5 次
它引用的顶会 Paper21
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng 等ICCV 2019 · 被引用 1,018 次
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao 等ICCV 2023 · 被引用 602 次
- NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving ScenarioTianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao 等AAAI 2024 · 被引用 314 次
- DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language ModelsLicheng Wen, Daocheng Fu, Xin Li, Xinyu Cai 等ICLR 2024 · 被引用 255 次
- PaLI: A Jointly-Scaled Multilingual Language-Image ModelXi Chen, Xiao Wang, Soravit Changpinyo, A. J. Piergiovanni 等ICLR 2023 · 被引用 194 次
相关 Paper
- LMDrive: Closed-Loop End-to-End Driving with Large Language ModelsHao Shao, Yuxuan Hu, Letian Wang, Guanglu Song 等CVPR 2024 · 被引用 114 次
- Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, EditingHao Fei, Shengqiong Wu, Hanwang Zhang, Tat-Seng Chua 等NeurIPS 2024 · 被引用 100 次
- Language-Image Models with 3D UnderstandingJang Hyun Cho, Boris Ivanovic, Yulong Cao, Edward Schmerling 等ICLR 2025 · 被引用 2 次
- DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video GenerationGuosheng Zhao, Xiaofeng Wang, Zheng Zhu, Xinze Chen 等AAAI 2025 · 被引用 31 次
- Generative Planning with 3D-Vision Language Pre-training for End-to-End Autonomous DrivingTengpeng Li, Hanli Wang, Xianfei Li, Wenlong Liao 等AAAI 2025 · 被引用 16 次
