MaskMentor: Unlocking the Potential of Masked Self-Teaching for Missing Modality RGB-D Semantic Segmentation
Zhida Zhao, Jia Li, Lijun Wang, Yifan Wang, Huchuan Lu
Abstract
Existing RGB-D semantic segmentation methods struggle to handle modality missing input, where only RGB images or depth maps are available, leading to degenerated segmentation performance. We tackle this issue using MaskMentor, a new pre-training framework for modality missing segmentation, which advances its counterparts via two novel designs: Masked Modality and Image Modeling (M2IM), and Self-Teaching via Token-Pixel Joint reconstruction (STTP). M2IM simulates modality missing scenarios by combining both modality- and patch-level random masking. Meanwhile, STTP offers an effective self-teaching strategy, where the trained network assumes a dual role, simultaneously acting as both the teacher and the student. The student with modality missing input is supervised by the teacher with complete modality input through both token- and pixel-wise masked modeling, closing the gap between missing and complete input modalities. By integrating M2IM and STTP, MaskMentor significantly improves the generalization ability of the trained model across diverse input conditions and outperforms state-of-the-art methods on two popular benchmarks by a considerable margin. Extensive ablation studies further verify the effectiveness of the above contributions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 859a6abf-90ce-4d83-ba86-a7f5dab0ca4aCited by top-tier papers7
- UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised CompensationXiaoqi Zhao, Youwei Pang, Chenyang Yu, Lihe Zhang et al.NeurIPS 2025 · 5 citations
- Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype DistillationJiaqi Tan, Xu Zheng, Yang LiuCVPR 2026 · 1 citation
- Mitigating Pervasive Modality Absence Through Multimodal Generalization and RefinementWuliang Huang, Yiqiang Chen, Xinlong Jiang, Chenlong Gao et al.AAAI 2025 · 1 citation
- Decoupled and Reusable Adaptation for Efficient Cross-Modal TransferYajing Liu, Yumeng Zhang, Yue Si, Baojie Fan et al.CVPR 2026
- Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic SegmentationJiaxin Cai, Jingze Su, Qi Li, Wenjie Yang et al.CVPR 2025
Related papers
- Multi-Modal Contrastive Masked Autoencoders: A Two-Stage Progressive Pre-training Approach for RGBD DatasetsMuhammad Abdullah Jamal, Omid MohareriCVPR 2025
- Mx2M: Masked Cross-Modality Modeling in Domain Adaptation for 3D Semantic SegmentationBoxiang Zhang, Zunran Wang, Yonggen Ling, Yuanyuan Guan et al.AAAI 2023 · 11 citations
- Tackling Dual-stage Missing Modalities in Brain Tumor Segmentation via Robust Modality Reconstruction and Prompt-guided Modality AdaptationYunpeng Zhao, Cheng Chen, Qing You Pang, Yibing Fu et al.AAAI 2026
- REDEEMing Modality Information Loss: Retrieval-Guided Conditional Generation for Severely Modality Missing LearningJian Lang, Rongpei Hong, Zhangtao Cheng, Ting Zhong et al.KDD 2025 · 5 citations
- DFormer: Rethinking RGBD Representation Learning for Semantic SegmentationBowen Yin, Xuying Zhang, Zhong-Yu Li, Li Liu et al.ICLR 2024 · 110 citations
