BEV-Guided Multi-Modality Fusion for Driving Perception
Yunze Man, Liang-Yan Gui, Yu-Xiong Wang
Abstract
Integrating multiple sensors and addressing diverse tasks in an end-to-end algorithm are challenging yet critical topics for autonomous driving. To this end, we introduce BEVGuide, a novel Bird's Eye-View (BEV) representation learning framework, representing the first attempt to unify a wide range of sensors under direct BEV guidance in an end-to-end fashion. Our architecture accepts input from a diverse sensor pool, including but not limited to Camera, Lidar and Radar sensors, and extracts BEV feature embeddings using a versatile and general transformer backbone. We design a BEV-guided multi-sensor attention block to take queries from BEV embeddings and learn the BEV representation from sensor-specific features. BEVGuide is efficient due to its lightweight backbone design and highly flexible as it supports almost any input sensor configurations. Extensive experiments demonstrate that our framework achieves exceptional performance in BEV perception tasks with a diverse sensor set. Project page is at https://yunzeman.github.io/BEVGuide .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21224416-c39b-47c5-9a2b-11a2179e02c9Cited by top-tier papers14
- Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene UnderstandingYunze Man, Shuhong Zheng, Zhipeng Bao, Martial Hebert et al.NeurIPS 2024 · 56 citations
- Situational Awareness Matters in 3D Vision Language ReasoningYunze Man, Liang-Yan Gui, Yu-Xiong WangCVPR 2024 · 9 citations
- SIRA: Scalable Inter-Frame Relation and Association for Radar PerceptionRyoma Yataka, Pu Wang, Petros Boufounos, Ryuhei TakahashiCVPR 2024 · 7 citations
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-SightYunze Man, Shihao Wang, Guowen Zhang, Johan Bjorck et al.CVPR 2026 · 6 citations
- NaMa: Neighbor-Aware Multi-Modal Adaptive Learning for Prostate Tumor Segmentation on Anisotropic MR ImagesRunqi Meng, Xiao Zhang, Shijie Huang, Yuning Gu et al.AAAI 2024 · 4 citations
Builds on9
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang et al.CVPR 2022 · 794 citations
- FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular CamerasAnthony Hu, Zak Murez, Nikhil Mohan, Sofía Dudas et al.ICCV 2021 · 329 citations
- Cross-view Transformers for real-time Map-view Semantic SegmentationBrady Zhou, Philipp KrähenbühlCVPR 2022 · 279 citations
- Structured Bird's-Eye-View Traffic Scene Understanding from Onboard ImagesYigit Baran Can, Alexander Liniger, Danda Pani Paudel, Luc Van GoolICCV 2021 · 147 citations
Related papers
- EVT: Efficient View Transformation for Multi-Modal 3D Object DetectionYongjin Lee, Hyeon Mun Jeong, Yurim Jeon, Sanghyun KimICCV 2025 · 5 citations
- UniDistill: A Universal Cross-Modality Knowledge Distillation Framework for 3D Object Detection in Bird's-Eye ViewShengchao Zhou, Weizhou Liu, Chen Hu, Shuchang Zhou et al.CVPR 2023
- DistillBEV: Boosting Multi-Camera 3D Object Detection with Cross-Modal Knowledge DistillationZeyu Wang, Dingwen Li, Chenxu Luo, Cihang Xie et al.ICCV 2023 · 65 citations
- UniFusion: Unified Multi-view Fusion Transformer for Spatial-Temporal Representation in Bird's-Eye-ViewZequn Qin, Jingyu Chen, Chao Chen, Xiaozhi Chen et al.ICCV 2023 · 38 citations
- Tri-Perspective View for Vision-Based 3D Semantic Occupancy PredictionYuanhui Huang, Wenzhao Zheng, Yunpeng Zhang, Jie Zhou et al.CVPR 2023
