FARFusion V2: A Geometry-based Radar-Camera Fusion Method on the Ground for Roadside Far-Range 3D Object Detection
Yao Li, Jiajun Deng, Yuxuan Xiao, Yingjie Wang, Xiaomeng Chu, Jianmin Ji, Yanyong Zhang
Abstract
Fusing the data of millimeter-wave Radar sensors and high-definition cameras has emerged as a viable approach to achieving precise 3D object detection for roadside traffic surveillance. For roadside perception systems, earlier studies have pointed out that it is better to perform the fusion on the 2D image plane than on the BEV plane (which is popular for on-car perception systems), especially when the perception range is large (e.g., >150m). Image-plane fusion requires critical transformations, like perspective projection from the Radar's BEV to the camera's 2D plane and reverse IPM. However, real-world issues like uneven terrain and sensor movement degrade these transformations' precision, impacting fusion effectiveness. To alleviate these issues, we propose a geometry-based Radar-camera fusion method on the ground, namely FARFusion V2. Specifically, we extend the ground-plane assumption in FARFusion[20] to support arbitrary shapes by formulating the ground height as an implicit representation based on geometric transformations. By incorporating the ground information, we can enhance Radar data with target height measurements. Consequently, we can thus project the enhanced Radar data onto the 2D plane to obtain more accurate depth information, thereby assisting the IPM process. A real-time parameterized transformation parameters estimation module is further introduced to refine the view transformation processes. Moreover, considering various measurement noises across these two sensors, we introduce an uncertainty-based depth fusion strategy into the 2D fusion process to maximize the probability of obtaining the optimal depth value. Extensive experiments are conducted on our collected roadside OWL benchmark, demonstrating the excellent localization capacity of FARFusion V2 in far-range scenarios. Our method achieves an average location accuracy of 0.771m when we extend the detection range up to 500m.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 66f350a0-edff-4e36-9345-9b4351ed3268Cited by top-tier papers1
Ask how each one uses itRelated papers
- RPGFusion: 4D Radar Prior-Guided Multi-Modal Fusion for 3D DetectionXin Qiu, Wenjie LiuCVPR 2026
- HGSFusion: Radar-Camera Fusion with Hybrid Generation and Synchronization for 3D Object DetectionZijian Gu, Jianwei Ma, Yan Huang, Honghao Wei et al.AAAI 2025 · 26 citations
- GSV2X: Geometry-Aware Uncertainty Modeling and Orthogonal Fusion for Robust Roadside PerceptionJianqiang Xu, Gensheng Pei, Huafeng Liu, Yazhou YaoCVPR 2026 · 2 citations
- RayFusion: Ray Fusion Enhanced Collaborative Visual PerceptionShaohong Wang, Lu Bin, Xinyu Xiao, Hanzhi Zhong et al.NeurIPS 2025 · 1 citation
- VIPS: real-time perception fusion for infrastructure-assisted autonomous drivingShuyao Shi, Jiahe Cui, Zhehao Jiang, Zhenyu Yan et al.MobiCom 2022 · 126 citations
