DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection
Xiang Li, Junbo Yin, Wei Li, Chengzhong Xu, Ruigang Yang, Jianbing Shen
摘要
Vehicle-to-Everything (V2X) collaborative perception has recently gained significant attention due to its capability to enhance scene understanding by integrating information from various agents, e.g., vehicles, and infrastructure. However, current works often treat the information from each agent equally, ignoring the inherent domain gap caused by the utilization of different LiDAR sensors of each agent, thus leading to suboptimal performance. In this paper, we propose DI-V2X, that aims to learn Domain-Invariant representations through a new distillation framework to mitigate the domain discrepancy in the context of V2X 3D object detection. DI-V2X comprises three essential components: a domain-mixing instance augmentation (DMA) module, a progressive domain-invariant distillation (PDD) module, and a domain-adaptive fusion (DAF) module. Specifically, DMA builds a domain-mixing 3D instance bank for the teacher and student models during training, resulting in aligned data representation. Next, PDD encourages the student models from different domains to gradually learn a domain-invariant feature representation towards the teacher, where the overlapping regions between agents are employed as guidance to facilitate the distillation process. Furthermore, DAF closes the domain gap between the students by incorporating calibration-aware domain-adaptive attention. Extensive experiments on the challenging DAIR-V2X and V2XSet benchmark datasets demonstrate DI-V2X achieves remarkable performance, outperforming all the previous V2X models. Code is available at https://github.com/Serenos/DI-V2X.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object DetectionJunbo Yin, Jianbing Shen, Runnan Chen, Wei Li 等CVPR 2024 · 被引用 73 次
- OLiDM: Object-aware LiDAR Diffusion Models for Autonomous DrivingTianyi Yan, Junbo Yin, Xianpeng Lang, Ruigang Yang 等AAAI 2025 · 被引用 16 次
- RoCo: Robust Cooperative Perception By Iterative Object Matching and Pose AdjustmentZhe Huang, Shuo Wang, Yongcai Wang, Wanting Li 等ACM MM 2024 · 被引用 12 次
- RCDN: Towards Robust Camera-Insensitivity Collaborative Perception via Dynamic Feature-based 3D Neural ModelingTianhang Wang, Fan Lu, Zehan Zheng, Zhijun Li 等NeurIPS 2024 · 被引用 11 次
- Redundant Queries in DETR-Based 3D Detection Methods: Unnecessary and PrunableLizhen Xu, Zehao Wu, Wenzhao Qiu, Shanmin Pang 等AAAI 2026 · 被引用 6 次
它引用的顶会 Paper7
- Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence MapsYue Hu, Shaoheng Fang, Zixing Lei, Yiqi Zhong 等NeurIPS 2022 · 被引用 537 次
- DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Cooperative 3D Object DetectionHaibao Yu, Yizhen Luo, Mao Shu, Yiyi Huo 等CVPR 2022 · 被引用 475 次
- Learning Distilled Collaboration Graph for Multi-Agent PerceptionYiming Li, Shunli Ren, Pengxiang Wu, Siheng Chen 等NeurIPS 2021 · 被引用 464 次
- V2X-Seq: A Large-Scale Sequential Dataset for Vehicle-Infrastructure Cooperative Perception and ForecastingHaibao Yu, Wenxian Yang, Hongzhi Ruan, Zhenwei Yang 等CVPR 2023
- When2com: Multi-Agent Perception via Communication Graph GroupingYen-Cheng Liu, Junjiao Tian, Nathaniel Glaser, Zsolt KiraCVPR 2020
相关 Paper
- DUSA: Decoupled Unsupervised Sim2Real Adaptation for Vehicle-to-Everything Collaborative PerceptionXianghao Kong, Wentao Jiang, Jinrang Jia, Yifeng Shi 等ACM MM 2023 · 被引用 18 次
- Privacy-Preserving V2X Collaborative Perception Integrating Unknown CollaboratorsBin Lu, Xinyu Xiao, Changzhou Zhang, Yang Zhou 等AAAI 2025
- UniDistill: A Universal Cross-Modality Knowledge Distillation Framework for 3D Object Detection in Bird's-Eye ViewShengchao Zhou, Weizhou Liu, Chen Hu, Shuchang Zhou 等CVPR 2023
- DSRC: Learning Density-Insensitive and Semantic-Aware Collaborative Representation Against CorruptionsJingyu Zhang, Yilei Wang, Lang Qian, Peng Sun 等AAAI 2025 · 被引用 13 次
- Leveraging Vision-Centric Multi-Modal Expertise for 3D Object DetectionLinyan Huang, Zhiqi Li, Chonghao Sima, Wenhai Wang 等NeurIPS 2023 · 被引用 26 次
