PrimKD: Primary Modality Guided Multimodal Fusion for RGB-D Semantic Segmentation
Zhiwei Hao, Zhongyu Xiao, Yong Luo, Jianyuan Guo, Jing Wang, Li Shen, Han Hu
摘要
The recent advancements in cross-modal transformers have demonstrated their superior performance in RGB-D segmentation tasks by effectively integrating information from both RGB and depth modalities. However, existing methods often overlook the varying levels of informative content present in each modality, treating them equally and using models of the same architecture. This oversight can potentially hinder segmentation performance, especially considering that RGB images typically contain significantly more information than depth images. To address this issue, we propose PrimKD, a knowledge distillation based approach that focuses on guided multimodal fusion, with an emphasis on leveraging the primary RGB modality. In our approach, we utilize a model trained exclusively on the RGB modality as the teacher, guiding the learning process of a student model that fuses both RGB and depth modalities. To prioritize information from the primary RGB modality while leveraging the depth modality, we incorporate primary focused feature reconstruction and a selective alignment scheme. This integration enhances the overall freature fusion, resulting in improved segmentation results. We evaluate our proposed method on the NYU Depth V2 and SUN-RGBD datasets, and the experimental results demonstrate the effectiveness of PrimKD. Specifically, our approach achieves mIoU scores of 57.8 and 52.5 on these two datasets, respectively, surpassing existing counterparts by 1.5 and 0.4 mIoU. The code is available at https://github.com/xiaoshideta/PrimKD.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic SegmentationBowen Yin, Jiao-Long Cao, Xuying Zhang, Yuming Chen 等NeurIPS 2025 · 被引用 8 次
- Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype DistillationJiaqi Tan, Xu Zheng, Yang LiuCVPR 2026 · 被引用 1 次
- LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural CracksHui Liu, Chen Jia, Fan Shi, Xu Cheng 等ACM MM 2025 · 被引用 1 次
- MixPrompt: Efficient Mixed Prompting for Multimodal Semantic SegmentationZhiwei Hao, Zhongyu Xiao, Jianyuan Guo, Li Shen 等NeurIPS 2025 · 被引用 1 次
- SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object DetectionJia Lin, Xiaofei Zhou, Jiyuan Liu, Runmin Cong 等AAAI 2026
相关 Paper
- G2D: Boosting Multimodal Learning with Gradient-Guided DistillationMohammed Rakib, Arunkumar BagavathiICCV 2025 · 被引用 1 次
- Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-ResolutionBaoli Sun, Xinchen Ye, Baopu Li, Haojie Li 等CVPR 2021
- SAM-Guided Semantic Knowledge Fusion for Visible-Infrared Object DetectionTing Li, Songtao Li, Shuaifeng Li, Xiaolin Qin 等ACM MM 2025 · 被引用 2 次
- Multimodal Decomposed Distillation with Instance Alignment and Uncertainty Compensation for Thermal Object DetectionYanfeng Liu, Lefei ZhangACM MM 2025 · 被引用 2 次
- C2KD: Bridging the Modality Gap for Cross-Modal Knowledge DistillationFushuo Huo, Wenchao Xu, Jingcai Guo, Haozhao Wang 等CVPR 2024
