Multi-View Attentive Contextualization for Multi-View 3D Object Detection
Xianpeng Liu, Ce Zheng, Ming Qian, Nan Xue, Chen Chen, Zhebin Zhang, Chen Li, Tianfu Wu
Abstract
We present Multi-View Attentive Contextualization (MvACon), a simple yet effective method for improving 2D-to-3D feature lifting in query-based multi-view 3D (MV3D) object detection. Despite remarkable progress witnessed in the field of query-based MV3D object detection, prior art often suffers from either the lack of exploiting high-resolution 2D features in dense attention-based lifting, due to high computational costs, or from insufficiently dense grounding of 3D queries to multi-scale 2D features in sparse attention-based lifting. Our proposed MvACon hits the two birds with one stone using a representationally dense yet computationally sparse attentive feature contextualization scheme that is agnostic to specific 2D-to-3D feature lifting approaches. In experiments, the proposed MvA-Con is thoroughly tested on the nuScenes benchmark, using both the BEVFormer and its recent 3D deformable attention (DFA3D) variant, as well as the PETR, showing consistent detection performance improvement, especially in enhancing performance in location, orientation, and velocity prediction. It is also tested on the Waymo-mini benchmark using BEVFormer with similar improvement. We qualitatively and quantitatively show that global cluster-based contexts effectively encode dense scene-level contexts for MV3D object detection. The promising results of our proposed MvA-Con reinforces the adage in computer vision - “(contextualized) feature matters”.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Towards Intrinsic-Aware Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Xiaoming LiuCVPR 2026 · 5 citations
- Unleashing the Temporal Potential of Stereo Event Cameras for Continuous-Time 3D Object DetectionJae-Young Kang, Hoonhee Cho, Kuk-Jin YoonICCV 2025 · 4 citations
- OcRFDet: Object-Centric Radiance Fields for Multi-View 3D Object Detection in Autonomous DrivingMingqian Ji, Shanshan Zhang, Jian YangICCV 2025 · 2 citations
- CHARM3R: Towards Unseen Camera Height Robust Monocular 3D DetectorAbhinav Kumar, Yuliang Guo, Zhihao Zhang, Xinyu Huang et al.ICCV 2025 · 1 citation
- Ev-3DOD: Pushing the Temporal Boundaries of 3D Object Detection with Event CamerasHoonhee Cho, Jae-Young Kang, Youngho Kim, Kuk-Jin YoonCVPR 2025
Builds on38
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo et al.NeurIPS 2021 · 2,148 citations
Related papers
- Object as Query: Lifting any 2D Object Detector to 3D DetectionZitian Wang, Zehao Huang, Jiahui Fu, Naiyan Wang et al.ICCV 2023 · 47 citations
- DFA3D: 3D Deformable Attention For 2D-to-3D Feature LiftingHongyang Li, Hao Zhang, Zhaoyang Zeng, Shilong Liu et al.ICCV 2023 · 40 citations
- Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object PredictionZhuofan Zong, Dongzhi Jiang, Guanglu Song, Zeyue Xue et al.ICCV 2023 · 63 citations
- SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera VideosHaisong Liu, Yao Teng, Tao Lu, Haiguang Wang et al.ICCV 2023 · 204 citations
- BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object DetectionZehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang et al.ICLR 2023 · 28 citations
