Distilling from Similar Tasks for Transfer Learning on a Budget
Kenneth Borup, Cheng Perng Phoo, Bharath Hariharan
Abstract
We address the challenge of getting efficient yet accurate recognition systems with limited labels. While recognition models improve with model size and amount of data, many specialized applications of computer vision have severe resource constraints both during training and inference. Transfer learning is an effective solution for training with few labels, however often at the expense of a computationally costly fine-tuning of large base models. We propose to mitigate this unpleasant trade-off between compute and accuracy via semi-supervised cross-domain distillation from a set of diverse source models. Initially, we show how to use task similarity metrics to select a single suitable source model to distill from, and that a good selection process is imperative for good downstream performance of a target model. We dub this approach DistillNearest. Though effective, DistillNearest assumes a single source model matches the target task, which is not always the case. To alleviate this, we propose a weighted multi-source distillation method to distill multiple source models trained on different domains weighted by their relevance for the target task into a single efficient model (named DistillWeighted). Our methods need no access to source data and merely need features and pseudo-labels of the source models. When the goal is accurate recognition under computational constraints, both DistillNearest and DistillWeighted approaches outperform both transfer learning from strong ImageNet initializations as well as state-of-the-art semisupervised techniques such as FixMatch. Averaged over 8 diverse target tasks our multi-source method outperforms the baselines by 5.6%-points and 4.5%-points, respectively. Code: github.com/Kennethborup/DistillWeighted
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cee9f5d-b8ab-4cd1-913e-104972663d26Cited by top-tier papers4
- Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific ModelsRaviteja Vemulapalli, Hadi Pouransari, Fartash Faghri, Sachin Mehta et al.ICML 2024 · 15 citations
- Scale-aware Recognition in Satellite Images under Resource ConstraintsShreelekha Revankar, Cheng Perng Phoo, Utkarsh Mall, Bharath Hariharan et al.ICLR 2025
- Objective drives the consistency of representational similarity across datasetsLaure Ciernik, Lorenz Linhardt, Marco Morik, Jonas Dippel et al.ICML 2025
- Learning Systems Expansion with Efficient Heterogeneity-aware Knowledge TransferGaole Dai, Huatao Xu, Yifan Yang, Rui Tan et al.AAAI 2026
Builds on16
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- Task2Vec: Task Embedding for Meta-LearningAlessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran et al.ICCV 2019 · 359 citations
Related papers
- Resource Efficient Domain AdaptationJunguang Jiang, Ximei Wang, Mingsheng Long, Jianmin WangACM MM 2020 · 26 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- Pay Attention to Features, Transfer Learn Faster CNNsKafeng Wang, Xitong Gao, Yiren Zhao, Xingjian Li et al.ICLR 2020 · 83 citations
- Probabilistic Model Distillation for Semantic CorrespondenceXin Li, Deng-Ping Fan, Fan Yang, Ao Luo et al.CVPR 2021
- Self-supervised Label Augmentation via Input TransformationsHankook Lee, Sung Ju Hwang, Jinwoo ShinICML 2020 · 218 citations
