Object as Query: Lifting any 2D Object Detector to 3D Detection
Zitian Wang, Zehao Huang, Jiahui Fu, Naiyan Wang, Si Liu
摘要
3D object detection from multi-view images has drawn much attention over the past few years. Existing methods mainly establish 3D representations from multi-view images and adopt a dense detection head for object detection, or employ object queries distributed in 3D space to localize objects. In this paper, we design Multi-View 2D Objects guided 3D Object Detector (MV2D), which can lift any 2D object detector to multi-view 3D object detection. Since 2D detections can provide valuable priors for object existence, MV2D exploits 2D detectors to generate object queries conditioned on the rich image semantics. These dynamically generated queries help MV2D to recall objects in the field of view and show a strong capability of localizing 3D objects. For the generated queries, we design a sparse cross attention module to force them to focus on the features of specific objects, which suppresses interference from noises. The evaluation results on the nuScenes dataset demonstrate the dynamic object queries and sparse feature aggregation can promote 3D detection capability. MV2D also exhibits a state-of-the-art performance among existing methods. We hope MV2D can serve as a new baseline for future research. Code is available at https://github.com/tusen-ai/MV2D .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- 3DPPE: 3D Point Positional Encoding for Transformer-based Multi-Camera 3D Object DetectionChangyong Shu, Jiajun Deng, Fisher Yu, Yifan LiuICCV 2023 · 被引用 34 次
- QE-BEV: Query Evolution for Bird's Eye View Object Detection in Varied ContextsJiawei Yao, Yingxin Lai, Hongrui Kou, Tong Wu 等ACM MM 2024 · 被引用 18 次
- RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric StrategiesXiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan 等ACM MM 2024 · 被引用 7 次
- Accelerate 3D Object Detection Models via Zero-Shot Attention Key PruningLizhen Xu, Xiuxiu Bai, Xiaojun Jia, Jianwu Fang 等ICCV 2025 · 被引用 1 次
- Dualad: Disentangling the Dynamic and Static World for End-to-End DrivingSimon Doll, Niklas Hanselmann, Lukas Schneider, Richard Schulz 等CVPR 2024
它引用的顶会 Paper12
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 被引用 542 次
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li 等ICCV 2023 · 被引用 513 次
相关 Paper
- Enhancing 3D Object Detection with 2D Detection-Guided Query AnchorsHaoxuanye Ji, Pengpeng Liang, Erkang ChengCVPR 2024
- Multi-View Attentive Contextualization for Multi-View 3D Object DetectionXianpeng Liu, Ce Zheng, Ming Qian, Nan Xue 等CVPR 2024 · 被引用 5 次
- Graph-DETR3D: Rethinking Overlapping Regions for Multi-View 3D Object DetectionZehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang 等ACM MM 2022 · 被引用 52 次
- STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object DetectionHuijie Fan, Pengrui Huang, Qiang Wang, Baojie Fan 等CVPR 2026
- SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera VideosHaisong Liu, Yao Teng, Tao Lu, Haiguang Wang 等ICCV 2023 · 被引用 204 次
