HAD: Heterogeneity-Aware Distillation for Lifelong Heterogeneous Learning
Xuerui Zhang, Xuehao Wang, Zhan Zhuang, Linglan Zhao, Ziyue Li, Xinmin Zhang, Zhihuan Song, Yu Zhang
Abstract
Lifelong learning aims to preserve knowledge acquired from previous tasks while incorporating knowledge from a sequence of new tasks. However, most prior work explores only streams of homogeneous tasks (e.g., only classification tasks) and neglects the scenario of learning across heterogeneous tasks that possess different structures of outputs. In this work, we formalize this broader setting as lifelong heterogeneous learning (LHL). Departing from conventional lifelong learning, the task sequence of LHL spans different task types, and the learner needs to retain heterogeneous knowledge for different output space structures. To instantiate the LHL, we focus on LHL in the context of dense prediction (LHL4DP), a realistic and challenging scenario. To this end, we propose the Heterogeneity-Aware Distillation (HAD) method, an exemplar-free approach that preserves previously gained heterogeneous knowledge by selfdistillation in each training phase. The proposed HAD comprises two complementary components, including a distribution-balanced heterogeneity-aware distillation loss to alleviate the global imbalance of prediction distribution and a salience-guided heterogeneity-aware distillation loss that concentrates learning on informative edge pixels extracted with the Sobel operator. Extensive experiments demonstrate that the proposed HAD method significantly outperforms existing methods in this new scenario.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on27
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Balanced MSE for Imbalanced Visual RegressionJiawei Ren, Mingyuan Zhang, Cunjun Yu, Ziwei LiuCVPR 2022 · 163 citations
- FeCAM: Exploiting the Heterogeneity of Class Distributions in Exemplar-Free Continual LearningDipam Goswami, Yuyang Liu, Bartlomiej Twardowski, Joost van de WeijerNeurIPS 2023 · 136 citations
Related papers
- Split-and-Bridge: Adaptable Class Incremental Learning within a Single Neural NetworkJong-Yeong Kim, Dong-Wan ChoiAAAI 2021 · 28 citations
- Lifelong Language Knowledge DistillationYung-Sung Chuang, Shang-Yu Su, Yun-Nung ChenEMNLP 2020 · 33 citations
- Multi-Domain Lifelong Visual Question Answering via Self-Critical DistillationMingrui Lao, Nan Pu, Yu Liu, Zhun Zhong et al.ACM MM 2023 · 5 citations
- Patch-based Knowledge Distillation for Lifelong Person Re-IdentificationZhicheng Sun, Yadong MuACM MM 2022 · 35 citations
- Distribution-Aware Knowledge Prototyping for Non-Exemplar Lifelong Person Re-IdentificationKunlun Xu, Xu Zou, Yuxin Peng, Jiahuan ZhouCVPR 2024 · 16 citations
