CAPE: Camera View Position Embedding for Multi-View 3D Object Detection
Kaixin Xiong, Shi Gong, Xiaoqing Ye, Xiao Tan, Ji Wan, Errui Ding, Jingdong Wang, Xiang Bai
摘要
In this paper, we address the problem of detecting 3D objects from multi-view images. Current query-based methods rely on global 3D position embeddings (PE) to learn the geometric correspondence between images and 3D space. We claim that directly interacting 2D image features with global 3D PE could increase the difficulty of learning view transformation due to the variation of camera extrinsics. Thus we propose a novel method based on CAmera view Position Embedding, called CAPE. We form the 3D position embeddings under the local camera-view coordinate system instead of the global coordinate system, such that 3D position embedding is free of encoding camera extrinsic parameters. Furthermore, we extend our CAPE to temporal modeling by exploiting the object queries of previous frames and encoding the ego motion for boosting 3D object detection. CAPE achieves the state-of-the-art performance (61.0% NDS and 52.5% mAP) among all LiDAR-free methods on nuScenes dataset. Codes and models are available. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal GenerationTianchen Deng, Xuefeng Chen, Yi Chen, Qu Chen 等CVPR 2026 · 被引用 31 次
- Query-based Temporal Fusion with Explicit Motion for 3D Object DetectionJinghua Hou, Zhe Liu, Dingkang Liang, Zhikang Zou 等NeurIPS 2023 · 被引用 28 次
- Leveraging Vision-Centric Multi-Modal Expertise for 3D Object DetectionLinyan Huang, Zhiqi Li, Chonghao Sima, Wenhai Wang 等NeurIPS 2023 · 被引用 26 次
- Dynamic Adapter Meets Prompt Tuning: Parameter-Efficient Transfer Learning for Point Cloud AnalysisXin Zhou, Dingkang Liang, Wei Xu, Xingkui Zhu 等CVPR 2024 · 被引用 25 次
- Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting LearningYang Jiao, Zequn Jie, Shaoxiang Chen, Lechao Cheng 等AAAI 2024 · 被引用 13 次
它引用的顶会 Paper24
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou 等AAAI 2021 · 被引用 1,128 次
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng 等ICCV 2021 · 被引用 974 次
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
相关 Paper
- Meta-Point Learning and Refining for Category-Agnostic Pose EstimationJunjie Chen, Jiebin Yan, Yuming Fang, Li NiuCVPR 2024
- Graph-DETR3D: Rethinking Overlapping Regions for Multi-View 3D Object DetectionZehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang 等ACM MM 2022 · 被引用 52 次
- Viewpoint Equivariance for Multi-View 3D Object DetectionDian Chen, Jie Li, Vitor Guizilini, Rares Ambrus 等CVPR 2023
- FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection TransformersHaisheng Su, Junjie Zhang, Feixiang Song, Sanping Zhou 等ICCV 2025
- 3DPPE: 3D Point Positional Encoding for Transformer-based Multi-Camera 3D Object DetectionChangyong Shu, Jiajun Deng, Fisher Yu, Yifan LiuICCV 2023 · 被引用 34 次
