RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusion
Xiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan, Houqiang Li, Yanyong Zhang
摘要
We propose Radar-Camera fusion transformer (RaC-Former) to boost the accuracy of 3D object detection by the following insight. The Radar-Camera fusion in outdoor 3D scene perception is capped by the image-to-BEV transformation-if the depth of pixels is not accurately estimated, the naive combination of BEV features actually integrates unaligned visual content. To avoid this problem, we propose a query-based framework that enables adaptive sampling of instance-relevant features from both the bird's-eye view (BEV) and the original image view. Furthermore, we enhance system performance by two key designs: optimizing query initialization and strengthening the representational capacity of BEV. For the former, we introduce an adaptive circular distribution in polar coordinates to refine the initialization of object queries, allowing for a distance-based adjustment of query density. For the latter, we initially incorporate a radar-guided depth head to refine the transformation from image view to BEV. Subsequently, we focus on leveraging the Doppler effect of radar and introduce an implicit dynamic catcher to capture the temporal elements within the BEV. Extensive experiments on nuScenes and View-of-Delft (VoD) datasets validate the merits of our design. Remarkably, our method achieves superior results of 64.9% mAP and 70.2% NDS on nuScenes. RaCFormer also secures the state-of-theart performance on the VoD dataset. Code is available at https://github.com/cxmomo/RaCFormer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- RaGS: Unleashing 3D Gaussian Splatting from 4D Radar and Monocular Cue for 3D Object DetectionXiaokai Bai, Chenxu Zhou, Lianqing Zheng, Jianan Liu 等CVPR 2026
- RPGFusion: 4D Radar Prior-Guided Multi-Modal Fusion for 3D DetectionXin Qiu, Wenjie LiuCVPR 2026
它引用的顶会 Paper21
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang 等CVPR 2022 · 被引用 794 次
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li 等ICCV 2023 · 被引用 513 次
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li 等ICCV 2021 · 被引用 404 次
- Unifying Voxel-based Representation with Transformer for 3D Object DetectionYanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li 等NeurIPS 2022 · 被引用 401 次
相关 Paper
- RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric StrategiesXiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan 等ACM MM 2024 · 被引用 7 次
- RobuRCDet: Enhancing Robustness of Radar-Camera Fusion in Bird's Eye View for 3D Object DetectionJingtong Yue, Zhiwei Lin, Xin Lin, Xiaoyu Zhou 等ICLR 2025
- RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object DetectionZhiwei Lin, Zhe Liu, Zhongyu Xia, Xinhao Wang 等CVPR 2024
- EVT: Efficient View Transformation for Multi-Modal 3D Object DetectionYongjin Lee, Hyeon Mun Jeong, Yurim Jeon, Sanghyun KimICCV 2025 · 被引用 5 次
- CRN: Camera Radar Net for Accurate, Robust, Efficient 3D PerceptionYoungseok Kim, Juyeb Shin, Sanmin Kim, In-Jae Lee 等ICCV 2023 · 被引用 134 次
