X-MoGe: A Cross-Modal Adaptation Framework with Mixture-of-Experts and Geometry Guidance for Heterogeneous Collaborative Perception
Wenkai Lin, Zhihong Liu, Chenglu Wen
Abstract
Multi-agent collaborative perception improves perception range and robustness in autonomous driving. However, most existing methods assume homogeneous sensors and perception networks, which is unrealistic in real-world heterogeneous systems. Differences in sensing modalities and independently trained models lead to significant semantic and geometric inconsistencies, limiting effective collaboration. To solve these problems, we propose a novel cross-modal adaptation framework with Mixture-of-Experts and geometry-guided fusion for heterogeneous collaborative perception, named X-MoGe. Specifically, we propose a Pixel-level Mixture-of-Experts (P-MoE) module, which adaptively models modality-specific semantic characteristics under heterogeneous sensing conditions. In addition, a geometry-guided feature fusion module incorporates explicit geometric priors to enforce spatial alignment and consistency in the BEV space. Extensive experiments on OPV2V and DAIR-V2X datasets demonstrate that the proposed method achieves state-of-the-art performance in heterogeneous collaborative perception.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 384212b1-dce2-4415-b6f3-fdc513a39dd0Builds on9
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Cooperative 3D Object DetectionHaibao Yu, Yizhen Luo, Mao Shu, Yiyi Huo et al.CVPR 2022 · 475 citations
- An Extensible Framework for Open Heterogeneous Collaborative PerceptionYifan Lu, Yue Hu, Yiqi Zhong, Dequan Wang et al.ICLR 2024 · 116 citations
- HM-ViT: Hetero-modal Vehicle-to-Vehicle Cooperative Perception with Vision TransformerHao Xiang, Runsheng Xu, Jiaqi MaICCV 2023 · 106 citations
Related papers
- Dr.Occ: Depth- and Region-Guided 3D Occupancy from Surround-View Cameras for Autonomous DrivingXubo Zhu, Haoyang Zhang, Fei He, Rui Wu et al.CVPR 2026 · 1 citation
- UniMM-V2X: MoE-Enhanced Multi-Level Fusion for End-to-End Cooperative Autonomous DrivingZiyi Song, Chen Xia, Chenbing Wang, Haibao Yu et al.AAAI 2026
- GT-Space: Enhancing Heterogeneous Collaborative Perception with Ground Truth Feature SpaceWentao Wang, Haoran Xu, Guang TanICLR 2026 · 2 citations
- STAMP: Scalable Task- And Model-agnostic Collaborative PerceptionXiangbo Gao, Runsheng Xu, Jiachen Li, Ziran Wang et al.ICLR 2025
- CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative PerceptionWeize Li, Yang Li, Quan Yuan, Xiaoyuan Fu et al.ICML 2026
