SA-BEV: Generating Semantic-Aware Bird's-Eye-View Feature for Multi-view 3D Object Detection
Jinqing Zhang, Yanan Zhang, Qingjie Liu, Yunhong Wang
摘要
Recently, the pure camera-based Bird's-Eye-View (BEV) perception provides a feasible solution for economical autonomous driving. However, the existing BEV-based multiview 3D detectors generally transform all image features into BEV features, without considering the problem that the large proportion of background information may submerge the object information. In this paper, we propose Semantic-Aware BEV Pooling (SA-BEVPool), which can filter out background information according to the semantic segmentation of image features and transform image features into semantic-aware BEV features. Accordingly, we propose BEV-Paste, an effective data augmentation strategy that closely matches with semantic-aware BEV feature. In addition, we design a Multi-Scale Cross-Task (MSCT) head, which combines task-specific and cross-task information to predict depth distribution and semantic segmentation more accurately, further improving the quality of semanticaware BEV feature. Finally, we integrate the above modules into a novel multi-view 3D object detection framework, namely SA-BEV. Experiments on nuScenes show that SA-BEV achieves state-of-the-art performance. Code has been available at https://github.com/mengtan00/SA-BEV.git .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- CRT-Fusion: Camera, Radar, Temporal Fusion Using Motion Information for 3D Object DetectionJisong Kim, Minjae Seong, Jun Won ChoiNeurIPS 2024 · 被引用 27 次
- GeoBEV: Learning Geometric BEV Representation for Multi-view 3D Object DetectionJinqing Zhang, Yanan Zhang, Yunlong Qi, Zehua Fu 等AAAI 2025 · 被引用 22 次
- BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-Based Roadside 3D Object DetectionWenjie Wang, Yehao Lu, Guangcong Zheng, Shuigen Zhan 等CVPR 2024 · 被引用 17 次
- PS-TTL: Prototype-based Soft-labels and Test-Time Learning for Few-shot Object DetectionYingjie Gao, Yanan Zhang, Ziyue Huang, Nanqing Liu 等ACM MM 2024 · 被引用 13 次
- Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting LearningYang Jiao, Zequn Jie, Shaoxiang Chen, Lechao Cheng 等AAAI 2024 · 被引用 13 次
它引用的顶会 Paper14
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou 等AAAI 2021 · 被引用 1,128 次
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li 等ICCV 2023 · 被引用 513 次
- Multimodal Virtual Point 3D DetectionTianwei Yin, Xingyi Zhou, Philipp KrähenbühlNeurIPS 2021 · 被引用 379 次
- CAT-Det: Contrastively Augmented Transformer for Multimodal 3D Object DetectionYanan Zhang, Jiaxin Chen, Di HuangCVPR 2022 · 被引用 138 次
相关 Paper
- MaskBEV: Towards A Unified Framework for BEV Detection and Map SegmentationXiao Zhao, Xukun Zhang, Dingkang Yang, Mingyang Sun 等ACM MM 2024 · 被引用 7 次
- SOGDet: Semantic-Occupancy Guided Multi-View 3D Object DetectionQiu Zhou, Jinming Cao, Hanchao Leng, Yifang Yin 等AAAI 2024 · 被引用 16 次
- Parametric Depth Based Feature Representation Learning for Object Detection and Segmentation in Bird's-Eye ViewJiayu Yang, Enze Xie, Miaomiao Liu, José M. ÁlvarezICCV 2023 · 被引用 9 次
- OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionZiteng Wen, Hai Xu, Chenyu Liu, Tao Guo 等ACM MM 2023 · 被引用 5 次
- BAEFormer: Bi-Directional and Early Interaction Transformers for Bird's Eye View Semantic SegmentationCong Pan, Yonghao He, Junran Peng, Qian Zhang 等CVPR 2023
