Distilling Knowledge from Heterogeneous Architectures for Semantic Segmentation
Yanglin Huang, Kai Hu, Yuan Zhang, Zhineng Chen, Xieping Gao
Abstract
Current knowledge distillation (KD) methods for semantic segmentation focus on guiding the student to imitate the teacher's knowledge within homogeneous architectures. However, these methods overlook the diverse knowledge contained in architectures with different inductive biases, which is crucial for enabling the student to acquire a more precise and comprehensive understanding of the data during distillation. To this end, we propose for the first time a generic knowledge distillation method for semantic segmentation from a heterogeneous perspective, named HeteroAKD. Due to the substantial disparities between heterogeneous architectures, such as CNN and Transformer, directly transferring cross-architecture knowledge presents significant challenges. To eliminate the influence of architecture-specific information, the intermediate features of both the teacher and student are skillfully projected into an aligned logits space. Furthermore, to utilize diverse knowledge from heterogeneous architectures and deliver customized knowledge required by the student, a teacher-student knowledge mixing mechanism (KMM) and a teacher-student knowledge evaluation mechanism (KEM) are introduced. These mechanisms are performed by assessing the reliability and its discrepancy between heterogeneous teacher-student knowledge. Extensive experiments conducted on three main-stream benchmarks using various teacher-student pairs demonstrate that our HeteroAKD framework outperforms state-of-the-art KD methods in facilitating distillation between heterogeneous architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dcd7edd5-afe1-4cc1-9de0-fdfccb2e32a9Cited by top-tier papers4
- Generalizable Knowledge Distillation from Vision Foundation Models for Semantic SegmentationChonghua Lv, Dong Zhao, Shuang Wang, Dou Quan et al.CVPR 2026 · 1 citation
- Heterogeneous Complementary DistillationLiuchi Xu, Hao Zheng, Lu Wang, Lisheng Xu et al.AAAI 2026
- Cross-Architecture Adaptation: Cloud-Edge Continual Test-Time Adaptation with Dynamic Sampling and Heterogeneous DistillationZirui Xu, Xianhang Chu, Jiahao Li, Xu Yang et al.CVPR 2026
- Progressive Multi-modal Knowledge Distillation for Multi-spectral Object Re-identificationAihua Zheng, Pengyu Li, Zi Wang, Jin TangAAAI 2026
Builds on14
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Channel-wise Knowledge Distillation for Dense Prediction*Changyong Shu, Yifan Liu, Jianfei Gao, Zheng Yan et al.ICCV 2021 · 432 citations
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and DepthThao Nguyen, Maithra Raghu, Simon KornblithICLR 2021 · 323 citations
- CycleMLP: A MLP-like Architecture for Dense PredictionShoufa Chen, Enze Xie, Chongjian Ge, Runjian Chen et al.ICLR 2022 · 254 citations
- Cross-Image Relational Knowledge Distillation for Semantic SegmentationChuanguang Yang, Helong Zhou, Zhulin An, Xue Jiang et al.CVPR 2022 · 228 citations
Related papers
- Fuse Before Transfer: Knowledge Fusion for Heterogeneous DistillationGuopeng Li, Qiang Wang, Ke Yan, Shouhong Ding et al.ICCV 2025 · 1 citation
- A Good Student is Cooperative and Reliable: CNN-Transformer Collaborative Learning for Semantic SegmentationJinjing Zhu, Yunhao Luo, Xu Zheng, Hao Wang et al.ICCV 2023 · 49 citations
- One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge DistillationZhiwei Hao, Jianyuan Guo, Kai Han, Yehui Tang et al.NeurIPS 2023 · 205 citations
- Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous DistillationJhe-Hao Lin, Yi Yao, Chan-Feng Hsu, Hong-Xia Xie et al.ICCV 2025 · 3 citations
- UniKD: Universal Knowledge Distillation for Mimicking Homogeneous or Heterogeneous Object DetectorsShanshan Lao, Guanglu Song, Boxiao Liu, Yu Liu et al.ICCV 2023 · 7 citations
