Co-Tuning for Transfer Learning
Kaichao You, Zhi Kou, Mingsheng Long, Jianmin Wang
Abstract
Fine-tuning pre-trained deep neural networks (DNNs) to a target dataset, also known as transfer learning, is widely used in computer vision and NLP. Because task-specific layers mainly contain categorical information and categories vary with datasets, practitioners only partially transfer pre-trained models by discarding task-specific layers and fine-tuning bottom layers. However, it is a reckless loss to simply discard task-specific parameters which take up as many as 20% of the total parameters in pre-trained models. To fully transfer pre-trained models, we propose a two-step framework named Co-Tuning: (i) learn the relationship between source categories and target categories from the pre-trained model with calibrated predictions; (ii) target labels (one-hot labels), as well as source labels (probabilistic labels) translated by the category relationship, collaboratively supervise the fine-tuning process. A simple instantiation of the framework shows strong empirical results in four visual classification tasks and one NLP classification task, bringing up to 20% relative improvement. While state-of-the-art fine-tuning techniques mainly focus on how to impose regularization when data are not abundant, Co-Tuning works not only in medium-scale datasets (100 samples per class) but also in large-scale datasets (1000 samples per class) where regularization-based methods bring no gains over the vanilla fine-tuning. Co-Tuning relies on a typically valid assumption that the pre-trained dataset is diverse enough, implying its broad application areas.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d7fa2cb-c5eb-4516-a56c-a59c1954d33bCited by top-tier papers24
- LogME: Practical Assessment of Pre-trained Models for Transfer LearningKaichao You, Yong Liu, Jianmin Wang, Mingsheng LongICML 2021 · 253 citations
- Self-Tuning for Data-Efficient Deep LearningXimei Wang, Jinghan Gao, Mingsheng Long, Jianmin WangICML 2021 · 79 citations
- Improved Fine-Tuning by Better Leveraging Pre-Training DataZiquan Liu, Yi Xu, Yuanhong Xu, Qi Qian et al.NeurIPS 2022 · 69 citations
- Zoo-Tuning: Adaptive Transfer from A Zoo of ModelsYang Shu, Zhi Kou, Zhangjie Cao, Jianmin Wang et al.ICML 2021 · 46 citations
- On the Connection between Pre-training Data Diversity and Fine-tuning RobustnessVivek Ramanujan, Thao Nguyen, Sewoong Oh, Ali Farhadi et al.NeurIPS 2023 · 40 citations
Builds on2
Related papers
- Fine-Tuning is Fine, if CalibratedZheda Mai, Arpita Chowdhury, Ping Zhang, Cheng-Hao Tu et al.NeurIPS 2024 · 34 citations
- Concept-wise Fine-tuning Matters in Preventing Negative TransferYunqiao Yang, Long-Kai Huang, Ying WeiICCV 2023 · 3 citations
- Co-Regularization Enhances Knowledge Transfer in High DimensionsShuo Shuo Liu, Haotian Lin, Matthew Reimherr, Runze LiNeurIPS 2025 · 2 citations
- AdaFilter: Adaptive Filter Fine-Tuning for Deep Transfer LearningYunhui Guo, Yandong Li, Liqiang Wang, Tajana RosingAAAI 2020 · 44 citations
- Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language ModelsZheng Zhao, Yftah Ziser, Shay B. CohenEMNLP 2024
