TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage Fusion
Yiran Wang, Jiaqi Li, Chaoyi Hong, Ruibo Li, Liusheng Sun, Xiao Song, Zhe Wang, Zhiguo Cao, Guosheng Lin
摘要
Radar-Camera depth estimation aims to predict dense and accurate metric depth by fusing input images and Radar data. Model efficiency is crucial for this task in pursuit of real-time processing on autonomous vehicles and robotic platforms. However, due to the sparsity of Radar returns, the prevailing methods adopt multi-stage frameworks with intermediate quasi-dense depth, which are time-consuming and not robust. To address these challenges, we propose TacoDepth, an efficient and accurate Radar-Camera depth estimation model with one-stage fusion. Specifically, the graph-based Radar structure extractor and the pyramidbased Radar fusion module are designed to capture and integrate the graph structures of Radar point clouds, delivering superior model efficiency and robustness without relying on the intermediate depth results. Moreover, TacoDepth can be flexible for different inference modes, providing a better balance of speed and accuracy. Extensive experiments are conducted to demonstrate the efficacy of our method. Compared with the previous state-of-the-art approach, TacoDepth improves depth accuracy and processing speed by 12.8% and 91.8%. Our work provides a new perspective on efficient Radar-Camera depth estimation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Radar-Guided Polynomial Fitting for Metric Depth EstimationPatrick Rim, Hyoungseob Park, Vadim Ezhov, Jeffrey Moon 等CVPR 2026 · 被引用 7 次
- Zero-Shot Depth Completion with Vision-Language ModelZhiqiang Yan, Yuan Wu, Gim Hee LeeCVPR 2026 · 被引用 1 次
- RPGFusion: 4D Radar Prior-Guided Multi-Modal Fusion for 3D DetectionXin Qiu, Wenjie LiuCVPR 2026
- Spectral-Geometric Neural Fields for Pose-Free LiDAR View SynthesisYinuo Jiang, Jun Cheng, Yiran Wang, Cheng ChengCVPR 2026
它引用的顶会 Paper22
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui 等ICCV 2019 · 被引用 3,193 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- CRN: Camera Radar Net for Accurate, Robust, Efficient 3D PerceptionYoungseok Kim, Juyeb Shin, Sanmin Kim, In-Jae Lee 等ICCV 2023 · 被引用 134 次
- R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object DetectionZhongyu Xia, Yousen Tang, Yongtao Wang, Zhifeng Wang 等CVPR 2026 · 被引用 4 次
- Echoes Beyond Points: Unleashing the Power of Raw Radar Data in Multi-modality FusionYang Liu, Feng Wang, Naiyan Wang, Zhaoxiang ZhangNeurIPS 2023 · 被引用 48 次
- CVFusion: Cross-View Fusion of 4D Radar and Camera for 3D Object DetectionHanzhi Zhong, Zhiyu Xiang, Ruoyu Xu, Jingyun Fu 等ICCV 2025 · 被引用 5 次
- Depth Estimation from Camera Image and mmWave Radar Point CloudAkash Deep Singh, Yunhao Ba, Ankur Sarker, Howard Zhang 等CVPR 2023
