Linking Modality Isolation in Heterogeneous Collaborative Perception
Changxing Liu, Zichen Chao, Siheng Chen
摘要
Collaborative perception leverages data exchange among multiple agents to enhance overall perception capabilities. However, heterogeneity across agents introduces domain gaps that hinder collaboration, and this is further exacerbated by an underexplored issue: modality isolation. It arises when multiple agents with different modalities never co-occur in any training data frame, enlarging cross-modal domain gaps. Existing alignment methods rely on supervision from spatially overlapping observations, thus fail to handle modality isolation. To address this challenge, we propose CodeAlign, the first efficient, co-occurrence-free alignment framework that smoothly aligns modalities via cross-modal feature-code-feature(FCF) translation. The key idea is to explicitly identify the representation consistency through codebook, and directly learn mappings between modality-specific feature spaces, thereby eliminating the need for spatial correspondence. Codebooks regularize feature spaces into code spaces, providing compact yet expressive representations. With a prepared code space for each modality, CodeAlign learns FCF translations that map features to the corresponding codes of other modalities, which are then decoded back into features in the target code space, enabling effective alignment. Experiments show that, when integrating three modalities, CodeAlign requires only 8% of the training parameters of prior alignment methods, reduces communication load by 1024x, and achieves state-of-the-art perception performance on both OPV2V and DAIR-V2X dataset. Code will be released on https://github.com/cxliu0314/CodeAlign.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence MapsYue Hu, Shaoheng Fang, Zixing Lei, Yiqi Zhong 等NeurIPS 2022 · 被引用 537 次
- DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Cooperative 3D Object DetectionHaibao Yu, Yizhen Luo, Mao Shu, Yiyi Huo 等CVPR 2022 · 被引用 475 次
- An Extensible Framework for Open Heterogeneous Collaborative PerceptionYifan Lu, Yue Hu, Yiqi Zhong, Dequan Wang 等ICLR 2024 · 被引用 116 次
- HM-ViT: Hetero-modal Vehicle-to-Vehicle Cooperative Perception with Vision TransformerHao Xiang, Runsheng Xu, Jiaqi MaICCV 2023 · 被引用 106 次
- Asynchrony-Robust Collaborative Perception via Bird's Eye View FlowSizhe Wei, Yuxi Wei, Yue Hu, Yifan Lu 等NeurIPS 2023 · 被引用 102 次
相关 Paper
- Communication-Efficient Collaborative Perception via Information Filling with CodebookYue Hu, Juntong Peng, Sifei Liu, Junhao Ge 等CVPR 2024
- Pragmatic Heterogeneous Collaborative Perception via Generative Communication MechanismJunfei Zhou, Penglin Dai, Quanmin Wei, Bingyi Liu 等NeurIPS 2025 · 被引用 11 次
- NegoCollab: A Common Representation Negotiation Approach for Heterogeneous Collaborative PerceptionCongzhang Shao, Quan Yuan, Guiyang Luo, Yue Hu 等NeurIPS 2025 · 被引用 7 次
- CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative PerceptionWeize Li, Yang Li, Quan Yuan, Xiaoyuan Fu 等ICML 2026
- GT-Space: Enhancing Heterogeneous Collaborative Perception with Ground Truth Feature SpaceWentao Wang, Haoran Xu, Guang TanICLR 2026 · 被引用 2 次
