Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype Distillation
Jiaqi Tan, Xu Zheng, Yang Liu
Abstract
Multimodal semantic segmentation (MMSS) faces significant challenges in real-world applications due to incomplete, degraded, or missing sensor data. To address this, we propose RobustSeg, an efficient teacher-student framework that enhances model robustness under missing-modality conditions while maintaining strong performance in fullmodality scenarios. RobustSeg adopts a feedback-based self-distillation paradigm consisting of two complementary stages. Firstly, we introduce Hybrid Prototype Distillation (HPD), which enables more reliable knowledge transfer of both cross-modal and modality-specific aspects. Concretely, combined with dominant-modality selection, HPD performs cross-modal semantic distillation with high-level semantic prototypes to reduce modality bias. Meanwhile, HPD conducts intra-class feature variation distillation for modality-specific structural details. Secondly, to enable the teacher model to gradually produce more balanced and robust modality representations, we make the student model provide feedback from the non-dominant modality to the teacher, benefiting the entire distillation process. Experiments on three datasets demonstrate that our method achieves state-of-the-art robustness (e.g., +2.40% missingmodality performance on DeLiVER) while causing almost no degradation in full-modality performance (only -0.1% mIoU). Moreover, evaluations using different backbones (AnySeg and CMNeXt) further validate the generalization ability of RobustSeg.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on25
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou et al.ICCV 2019 · 625 citations
- Rethinking Semantic Segmentation: A Prototype ViewTianfei Zhou, Wenguan Wang, Ender Konukoglu, Luc Van GoolCVPR 2022 · 353 citations
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
- Multimodal Token Fusion for Vision TransformersYikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang et al.CVPR 2022 · 214 citations
Related papers
- A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing ModalitiesMingcheng Li, Dingkang Yang, Yuxuan Lei, Shunli Wang et al.AAAI 2024 · 71 citations
- Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete ModalitiesMingcheng Li, Dingkang Yang, Xiao Zhao, Shuaibing Wang et al.CVPR 2024
- MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge DistillationHui Li, Pengfei Yang, Juanyang Chen, Le Dong et al.ACM MM 2025 · 4 citations
- C2KD: Bridging the Modality Gap for Cross-Modal Knowledge DistillationFushuo Huo, Wenchao Xu, Jingcai Guo, Haozhao Wang et al.CVPR 2024
- Distilling Knowledge from Heterogeneous Architectures for Semantic SegmentationYanglin Huang, Kai Hu, Yuan Zhang, Zhineng Chen et al.AAAI 2025 · 4 citations
