Multi-Modal Fusion Transformer for End-to-End Autonomous Driving
Aditya Prakash, Kashyap Chitta, Andreas Geiger
Abstract
How should representations from complementary sensors be integrated for autonomous driving? Geometrybased sensor fusion has shown great promise for perception tasks such as object detection and motion forecasting. However, for the actual driving task, the global context of the 3D scene is key, e.g. a change in traffic light state can affect the behavior of a vehicle geometrically distant from that traffic light. Geometry alone may therefore be insufficient for effectively fusing representations in end-to-end driving models. In this work, we demonstrate that imitation learning policies based on existing sensor fusion methods under-perform in the presence of a high density of dynamic agents and complex scenarios, which require global contextual reasoning, such as handling traffic oncoming from multiple directions at uncontrolled intersections. Therefore, we propose TransFuser, a novel Multi-Modal Fusion Transformer, to integrate image and LiDAR representations using attention. We experimentally validate the efficacy of our approach in urban settings involving complex scenarios using the CARLA urban driving simulator. Our approach achieves state-of-the-art driving performance while reducing collisions by 76% compared to geometry-based fusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59c4ec9a-2624-4747-b5e4-f622b44d2845Cited by top-tier papers116
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao et al.ICCV 2023 · 602 citations
- Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong BaselinePenghao Wu, Xiaosong Jia, Li Chen, Junchi Yan et al.NeurIPS 2022 · 444 citations
- TransWeather: Transformer-based Restoration of Images Degraded by Adverse Weather ConditionsJeya Maria Jose Valanarasu, Rajeev Yasarla, Vishal M. PatelCVPR 2022 · 350 citations
- NEAT: Neural Attention Fields for End-to-End Autonomous DrivingKashyap Chitta, Aditya Prakash, Andreas GeigerICCV 2021 · 274 citations
- VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic PlanningBo Jiang, Shaoyu Chen, Hao Gao, Bencheng Liao et al.ICLR 2026 · 259 citations
Builds on17
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Exploring the Limitations of Behavior Cloning for Autonomous DrivingFelipe Codevilla, Eder Santana, Antonio M. López, Adrien GaidonICCV 2019 · 666 citations
- STGAT: Modeling Spatial-Temporal Interactions for Human Trajectory PredictionYingfan Huang, Huikun Bi, Zhaoxin Li, Tianlu Mao et al.ICCV 2019 · 615 citations
- The Trajectron: Probabilistic Multi-Agent Trajectory Modeling With Dynamic Spatiotemporal GraphsBoris Ivanovic, Marco PavoneICCV 2019 · 473 citations
Related papers
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang et al.CVPR 2022 · 794 citations
- DiffE2E: Rethinking End-to-End Driving with a Hybrid Diffusion-Regression-Classification PolicyRui Zhao, Yuze Fan, Ziguo Chen, Fei Gao et al.NeurIPS 2025 · 7 citations
- LEAD: Minimizing Learner-Expert Asymmetry in End-to-End DrivingLong Nguyen, Micha Fauth, Bernhard Jaeger, Daniel Dauner et al.CVPR 2026 · 28 citations
- LIFT: Learning 4D LiDAR Image Fusion Transformer for 3D Object DetectionYihan Zeng, Da Zhang, Chunwei Wang, Zhenwei Miao et al.CVPR 2022 · 36 citations
- GaussianFusion: Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous DrivingShuai Liu, Quanmin Liang, Zefeng Li, Boyang Li et al.NeurIPS 2025 · 20 citations
