Flow-Based Feature Fusion for Vehicle-Infrastructure Cooperative 3D Object Detection
Haibao Yu, Yingjuan Tang, Enze Xie, Jilei Mao, Ping Luo, Zaiqing Nie
摘要
Cooperatively utilizing both ego-vehicle and infrastructure sensor data can significantly enhance autonomous driving perception abilities. However, the uncertain temporal asynchrony and limited communication conditions can lead to fusion misalignment and constrain the exploitation of infrastructure data. To address these issues in vehicle-infrastructure cooperative 3D (VIC3D) object detection, we propose the Feature Flow Net (FFNet), a novel cooperative detection framework. FFNet is a flow-based feature fusion framework that uses a feature flow prediction module to predict future features and compensate for asynchrony. Instead of transmitting feature maps extracted from still-images, FFNet transmits feature flow, leveraging the temporal coherence of sequential infrastructure frames. Furthermore, we introduce a self-supervised training approach that enables FFNet to generate feature flow with feature prediction ability from raw infrastructure sequences. Experimental results demonstrate that our proposed method outperforms existing cooperative detection methods while only requiring about 1/100 of the transmission cost of raw data and covers all latency in one model on the DAIR-V2X dataset. The code is available at https://github.com/haibao-yu/FFNet-VIC3D.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- End-to-End Autonomous Driving Through V2X CooperationHaibao Yu, Wenxian Yang, Jiaru Zhong, Zhenwei Yang 等AAAI 2025 · 被引用 56 次
- Learning Cooperative Trajectory Representations for Motion ForecastingHongzhi Ruan, Haibao Yu, Wenxian Yang, Siqi Fan 等NeurIPS 2024 · 被引用 36 次
- V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and PredictionZewei Zhou, Hao Xiang, Zhaoliang Zheng, Seth Z. Zhao 等ICCV 2025 · 被引用 15 次
- CoST: Efficient Collaborative Perception from Unified Spatiotemporal PerspectiveZongheng Tang, Yi Liu, Yifan Sun, Yulu Gao 等ICCV 2025 · 被引用 6 次
- INSTINCT: Instance-Level Interaction Architecture for Query-Based Collaborative PerceptionYunjiang Xu, Lingzhi Li, Jin Wang, Yupeng Ouyang 等ICCV 2025 · 被引用 5 次
它引用的顶会 Paper11
- STD: Sparse-to-Dense 3D Object Detector for Point CloudZetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen 等ICCV 2019 · 被引用 840 次
- Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence MapsYue Hu, Shaoheng Fang, Zixing Lei, Yiqi Zhong 等NeurIPS 2022 · 被引用 537 次
- DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Cooperative 3D Object DetectionHaibao Yu, Yizhen Luo, Mao Shu, Yiyi Huo 等CVPR 2022 · 被引用 475 次
- Learning Distilled Collaboration Graph for Multi-Agent PerceptionYiming Li, Shunli Ren, Pengxiang Wu, Siheng Chen 等NeurIPS 2021 · 被引用 464 次
- Rope3D: The Roadside Perception Dataset for Autonomous Driving and Monocular 3D Object Detection TaskXiaoqing Ye, Mao Shu, Hanyu Li, Yifeng Shi 等CVPR 2022 · 被引用 130 次
相关 Paper
- TraF-Align: Trajectory-aware Feature Alignment for Asynchronous Multi-agent PerceptionZhiying Song, Lei Yang, Fuxi Wen, Jun LiCVPR 2025
- BEVSync: Asynchronous Data Alignment for Camera-based Vehicle-Infrastructure Cooperative Perception Under Uncertain DelaysWentao Wang, Jiaqian Wang, Yuxin Deng, Guang TanAAAI 2025 · 被引用 2 次
- ViTraj: Learning Dual-Side Representations for Vehicle-Infrastructure Cooperative Trajectory PredictionShengzhe You, Libo Weng, Fei GaoACM MM 2025
- SparseAlign: a Fully Sparse Framework for Cooperative Object DetectionYunshuang Yuan, Yan Xia, Daniel Cremers, Monika SesterCVPR 2025
- Asynchrony-Robust Collaborative Perception via Bird's Eye View FlowSizhe Wei, Yuxi Wei, Yue Hu, Yifan Lu 等NeurIPS 2023 · 被引用 102 次
