MaskBEV: Towards A Unified Framework for BEV Detection and Map Segmentation
Xiao Zhao, Xukun Zhang, Dingkang Yang, Mingyang Sun, Mingcheng Li, Shunli Wang, Lihua Zhang
Abstract
Accurate and robust multimodal multi-task perception is crucial for modern autonomous driving systems. However, current multimodal perception research follows independent paradigms designed for specific perception tasks, leading to a lack of complementary learning among tasks and decreased performance in multi-task learning (MTL) due to joint training. In this paper, we propose MaskBEV, a masked attention-based MTL paradigm that unifies 3D object detection and bird's eye view (BEV) map segmentation. MaskBEV introduces a task-agnostic Transformer decoder to process these diverse tasks, enabling MTL to be completed in a unified decoder without requiring additional design of specific task heads. To fully exploit the complementary information between BEV map segmentation and 3D object detection tasks in BEV space, we propose spatial modulation and scene-level context aggregation strategies. These strategies consider the inherent dependencies between BEV segmentation and 3D detection, naturally boosting MTL performance. Extensive experiments on nuScenes dataset show that compared with previous state-of-the-art MTL methods, MaskBEV achieves 1.3 NDS improvement in 3D object detection and 2.7 mIoU improvement in BEV map segmentation, while also demonstrating slightly leading inference speed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7494c411-de8e-498c-b487-bd35df93ee5aCited by top-tier papers3
- RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy PredictionBaojie Fan, Xiaotian Li, Yuhan Zhou, Yuyu Jiang et al.ICCV 2025 · 1 citation
- Adaptive-Smooth LiDAR-Camera Knowledge Distillation with Heterogeneous Fusion for Multi-View 3D Object DetectionRui Zhao, Shuoyao Wang, Xinhu Zheng, Shijian GaoAAAI 2026
- MAESTRO: Task-Relevant Optimization Via Adaptive Feature Enhancement and Suppression for Multi-Task 3D PerceptionChangwon Kang, Jisong Kim, Hongjae Shin, Junseo Park et al.ICCV 2025
Builds on33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
Related papers
- SA-BEV: Generating Semantic-Aware Bird's-Eye-View Feature for Multi-view 3D Object DetectionJinqing Zhang, Yanan Zhang, Qingjie Liu, Yunhong WangICCV 2023 · 41 citations
- MTA: Multimodal Task Alignment for BEV Perception and CaptioningYunsheng Ma, Burhan Yaman, Xin Ye, Jingru Luo et al.CVPR 2026
- A Versatile Multi-View Framework for LiDAR-based 3D Object Detection with Guidance from Panoptic SegmentationHamidreza Fazlali, Yixuan Xu, Yuan Ren, Bingbing LiuCVPR 2022 · 23 citations
- UniTR: A Unified and Efficient Multi-Modal Transformer for Bird's-Eye-View RepresentationHaiyang Wang, Hao Tang, Shaoshuai Shi, Aoxue Li et al.ICCV 2023 · 106 citations
- MetaBEV: Solving Sensor Failures for 3D Detection and Map SegmentationChongjian Ge, Junsong Chen, Enze Xie, Zhongdao Wang et al.ICCV 2023 · 64 citations
