Distilling from Similar Tasks for Transfer Learning on a Budget
Kenneth Borup, Cheng Perng Phoo, Bharath Hariharan
摘要
We address the challenge of getting efficient yet accurate recognition systems with limited labels. While recognition models improve with model size and amount of data, many specialized applications of computer vision have severe resource constraints both during training and inference. Transfer learning is an effective solution for training with few labels, however often at the expense of a computationally costly fine-tuning of large base models. We propose to mitigate this unpleasant trade-off between compute and accuracy via semi-supervised cross-domain distillation from a set of diverse source models. Initially, we show how to use task similarity metrics to select a single suitable source model to distill from, and that a good selection process is imperative for good downstream performance of a target model. We dub this approach DistillNearest. Though effective, DistillNearest assumes a single source model matches the target task, which is not always the case. To alleviate this, we propose a weighted multi-source distillation method to distill multiple source models trained on different domains weighted by their relevance for the target task into a single efficient model (named DistillWeighted). Our methods need no access to source data and merely need features and pseudo-labels of the source models. When the goal is accurate recognition under computational constraints, both DistillNearest and DistillWeighted approaches outperform both transfer learning from strong ImageNet initializations as well as state-of-the-art semisupervised techniques such as FixMatch. Averaged over 8 diverse target tasks our multi-source method outperforms the baselines by 5.6%-points and 4.5%-points, respectively. Code: github.com/Kennethborup/DistillWeighted
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific ModelsRaviteja Vemulapalli, Hadi Pouransari, Fartash Faghri, Sachin Mehta 等ICML 2024 · 被引用 15 次
- Scale-aware Recognition in Satellite Images under Resource ConstraintsShreelekha Revankar, Cheng Perng Phoo, Utkarsh Mall, Bharath Hariharan 等ICLR 2025
- Objective drives the consistency of representational similarity across datasetsLaure Ciernik, Lorenz Linhardt, Marco Morik, Jonas Dippel 等ICML 2025
- Learning Systems Expansion with Efficient Heterogeneity-aware Knowledge TransferGaole Dai, Huatao Xu, Yifan Yang, Rui Tan 等AAAI 2026
它引用的顶会 Paper16
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- Task2Vec: Task Embedding for Meta-LearningAlessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran 等ICCV 2019 · 被引用 359 次
相关 Paper
- Resource Efficient Domain AdaptationJunguang Jiang, Ximei Wang, Mingsheng Long, Jianmin WangACM MM 2020 · 被引用 26 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
- Pay Attention to Features, Transfer Learn Faster CNNsKafeng Wang, Xitong Gao, Yiren Zhao, Xingjian Li 等ICLR 2020 · 被引用 83 次
- Probabilistic Model Distillation for Semantic CorrespondenceXin Li, Deng-Ping Fan, Fan Yang, Ao Luo 等CVPR 2021
- Self-supervised Label Augmentation via Input TransformationsHankook Lee, Sung Ju Hwang, Jinwoo ShinICML 2020 · 被引用 218 次
