ObjectFusion: Multi-modal 3D Object Detection with Object-Centric Fusion
Qi Cai, Yingwei Pan, Ting Yao, Chong-Wah Ngo, Tao Mei
摘要
Recent progress on multi-modal 3D object detection has featured BEV (Bird-Eye-View) based fusion, which effectively unifies both LiDAR point clouds and camera images in a shared BEV space. Nevertheless, it is not trivial to perform camera-to-BEV transformation due to the inherently ambiguous depth estimation of each pixel, resulting in spatial misalignment between these two multi-modal features. Moreover, such transformation also inevitably leads to projection distortion of camera image features in BEV space. In this paper, we propose a novel Object-centric Fusion (ObjectFusion) paradigm, which completely gets rid of camera-to-BEV transformation during fusion to align object-centric features across different modalities for 3D object detection. ObjectFusion first learns three kinds of modality-specific feature maps (i.e., voxel, BEV, and image features) from LiDAR point clouds and its BEV projections, camera images. Then a set of 3D object proposals are produced from the BEV features via a heatmap-based proposal generator. Next, the 3D object proposals are reprojected back to voxel, BEV, and image spaces. We leverage voxel and RoI pooling to generate spatially aligned object-centric features for each modality. All the object-centric features of three modalities are further fused at object level, which is finally fed into the detection heads. Extensive experiments on nuScenes dataset demonstrate the superiority of our Ob-jectFusion, by achieving 69.8% mAP on nuScenes validation set and improving BEVFusion by 1.3%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving ScenesXiaoyu Zhou, Zhiwei Lin, Xiaojun Shan, Yongtao Wang 等CVPR 2024 · 被引用 166 次
- IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object DetectionJunbo Yin, Jianbing Shen, Runnan Chen, Wei Li 等CVPR 2024 · 被引用 73 次
- DrivingForward: Feed-forward 3D Gaussian Splatting for Driving Scene Reconstruction from Flexible Surround-view InputQijian Tian, Xin Tan, Yuan Xie, Lizhuang MaAAAI 2025 · 被引用 45 次
- UniPAD: A Universal Pre-Training Paradigm for Autonomous DrivingHonghui Yang, Sha Zhang, Di Huang, Xiaoyang Wu 等CVPR 2024 · 被引用 31 次
- Height-Fidelity Dense Global Fusion for Multi-Modal 3D Object DetectionHanshi Wang, Jin Gao, Weiming Hu, Zhipeng ZhangICCV 2025 · 被引用 9 次
它引用的顶会 Paper27
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou 等AAAI 2021 · 被引用 1,128 次
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- STD: Sparse-to-Dense 3D Object Detector for Point CloudZetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen 等ICCV 2019 · 被引用 840 次
相关 Paper
- GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object DetectionXiaotian Li, Baojie Fan, Jiandong Tian, Huijie FanCVPR 2024
- BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object DetectionGuowen Zhang, Chenhang He, Liyi Chen, Lei ZhangAAAI 2026 · 被引用 2 次
- DLFusion: Painting-Depth Augmenting-LiDAR for Multimodal Fusion 3D Object DetectionJunyin Wang, Chenghu Du, Hui Li, Shengwu XiongACM MM 2023 · 被引用 3 次
- BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object DetectionZehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang 等ICLR 2023 · 被引用 28 次
- EVT: Efficient View Transformation for Multi-Modal 3D Object DetectionYongjin Lee, Hyeon Mun Jeong, Yurim Jeon, Sanghyun KimICCV 2025 · 被引用 5 次
