Privacy-Preserving V2X Collaborative Perception Integrating Unknown Collaborators
Bin Lu, Xinyu Xiao, Changzhou Zhang, Yang Zhou, Zhiyu Xiang, Hangguan Shan, Eryun Liu
Abstract
Vehicle-to-everything (V2X) collaborative perception has recently gained increasing attention in autonomous driving due to its ability to enhance scene understanding by integrating information from other collaborators, e.g. vehicles or infrastructure. Existing algorithms usually share deep features to achieve a trade-off between accuracy and bandwidth. However, most of these methods require joint training of all agents, which results in privacy leakage and is impractical and unacceptable in the real world. Sharing prediction results seems to be a direct solution, but its performance is suboptimal and sensitive to localization noise and communication delay. In this paper, we propose a privacy-preserving collaborative perception framework, where each agent is separately trained with its own dataset and the ego vehicle needs to integrate with completely unknown collaborators. Specifically, we propose MSD, a multi-scale feature fusion method combined with deformable attention, to better fuse features of different agents. We also propose a plug-in domain adapter to align the features from unknown collaborators to ego-domain. Extensive experiments on the challenging DAIR-V2X and V2V4Real demonstrate that: 1) MSD achieves remarkable performance, outperforming others by at least 2.8% and 6.7% in AP0.7 on DAIR-V2X and V2V4Real, respectively; 2) After domain adaptation, it significantly outperforms the No Fusion, Late Fusion scenarios and can approach or even surpass the performance of joint training. We truly achieves privacy-preserving collaboration, providing a new paradigm for the study of collaborative perception, which is crucial for practical applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- STD: Sparse-to-Dense 3D Object Detector for Point CloudZetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen et al.ICCV 2019 · 840 citations
- Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence MapsYue Hu, Shaoheng Fang, Zixing Lei, Yiqi Zhong et al.NeurIPS 2022 · 537 citations
Related papers
- DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object DetectionXiang Li, Junbo Yin, Wei Li, Chengzhong Xu et al.AAAI 2024 · 56 citations
- DUSA: Decoupled Unsupervised Sim2Real Adaptation for Vehicle-to-Everything Collaborative PerceptionXianghao Kong, Wentao Jiang, Jinrang Jia, Yifeng Shi et al.ACM MM 2023 · 18 citations
- Pragmatic Heterogeneous Collaborative Perception via Generative Communication MechanismJunfei Zhou, Penglin Dai, Quanmin Wei, Bingyi Liu et al.NeurIPS 2025 · 11 citations
- What2comm: Towards Communication-efficient Collaborative Perception via Feature DecouplingKun Yang, Dingkang Yang, Jingyu Zhang, Hanqi Wang et al.ACM MM 2023 · 58 citations
- TransIFF: An Instance-Level Feature Fusion Framework for Vehicle-Infrastructure Cooperative 3D Detection with TransformersZiming Chen, Yifeng Shi, Jinrang JiaICCV 2023 · 55 citations
