Resilient Sensor Fusion Under Adverse Sensor Failures via Multi-Modal Expert Fusion
Konyul Park, Yecheol Kim, Daehun Kim, Jun Won Choi
摘要
Modern autonomous driving perception systems utilize complementary multi-modal sensors, such as LiDAR and cameras. Although sensor fusion architectures enhance performance in challenging environments, they still suffer significant performance drops under severe sensor failures, such as LiDAR beam reduction, LiDAR drop, limited field of view, camera drop, and occlusion. This limitation stems from inter-modality dependencies in current sensor fusion frameworks. In this study, we introduce an efficient and robust LiDAR-camera 3D object detector, referred to as MoME, which can achieve robust performance through a mixture of experts approach. Our MoME fully decouples modality dependencies using three parallel expert decoders, which use camera features, LiDAR features, or a combination of both to decode object queries, respectively. We propose Multi-Expert Decoding (MED) framework, where each query is decoded selectively using one of three expert decoders. MoME utilizes an Adaptive Query Router (AQR) to select the most appropriate expert decoder for each query based on the quality of camera and LiDAR features. This ensures that each query is processed by the best-suited expert, resulting in robust performance across diverse sensor failure scenarios. We evaluated the performance of MoME on the nuScenes-R benchmark. Our MoME achieved state-of-the-art performance in extreme weather and sensor failure conditions, significantly outperforming the existing models across various sensor failure scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object DetectionYuchen Wu, Kun Wang, Yining Pan, Na ZhaoCVPR 2026 · 被引用 4 次
- AnyMod-LLVE: Low-Light Video Enhancement with Modality-Agnostic InferenceHangfeng Liang, Yutao Hu, Yanhan Hu, Xiaohan Wu 等ICML 2026
- NeuroMamba: A Universal Spatiotemporal Module for Robust Perception in Degraded Sensory StreamsJinfeng Li, Huijia Song, Xiangyue Hu, HanLiang Zhou 等ICML 2026
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang 等CVPR 2022 · 被引用 794 次
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkTingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia 等NeurIPS 2022 · 被引用 762 次
- DeepInteraction: 3D Object Detection via Modality InteractionZeyu Yang, Jiaqi Chen, Zhenwei Miao, Wei Li 等NeurIPS 2022 · 被引用 268 次
相关 Paper
- MetaBEV: Solving Sensor Failures for 3D Detection and Map SegmentationChongjian Ge, Junsong Chen, Enze Xie, Zhongdao Wang 等ICCV 2023 · 被引用 64 次
- L4DR: LiDAR-4DRadar Fusion for Weather-Robust 3D Object DetectionXun Huang, Ziyu Xu, Hai Wu, Jinlong Wang 等AAAI 2025 · 被引用 39 次
- Taming Cascaded Mixture-of-Experts for Modality-missing Multi-modal Salient Object DetectionKunpeng Wang, Feifan Sun, Keke ChenAAAI 2026
- Benchmarking Robustness of 3D Object Detection to Common Corruptions in Autonomous DrivingYinpeng Dong, Caixin Kang, Jinlai Zhang, Zijian Zhu 等CVPR 2023
- MSeg3D: Multi-Modal 3D Semantic Segmentation for Autonomous DrivingJiale Li, Hang Dai, Hao Han, Yong DingCVPR 2023
