Concept-wise Fine-tuning Matters in Preventing Negative Transfer
Yunqiao Yang, Long-Kai Huang, Ying Wei
Abstract
A multitude of prevalent pre-trained models mark a major milestone in the development of artificial intelligence, while fine-tuning has been a common practice that enables pretrained models to figure prominently in a wide array of target datasets. Our empirical results reveal that off-the-shelf finetuning techniques are far from adequate to mitigate negative transfer caused by two types of underperforming features in a pre-trained model, including rare features and spuriously correlated features. Rooted in structural causal models of predictions after fine-tuning, we propose a Concept-wise fine-tuning (Concept-Tuning) approach which refines feature representations in the level of patches with each patch encoding a concept. Concept-Tuning minimizes the negative impacts of rare features and spuriously correlated features by (1) maximizing the mutual information between examples in the same category with regard to a slice of rare features (a patch) and (2) applying front-door adjustment via attention neural networks in channels and feature slices (patches). The proposed Concept-Tuning consistently and significantly (by up to 4.76%) improves prior state-of-the-art fine-tuning methods on eleven datasets, diverse pre-training strategies (supervised and self-supervised ones), various network architectures, and sample sizes in a target dataset. * Part of the work was done when the author interned at Tencent AI Lab. † Corresponding Author (a) Train from scratch (c) Fine-tuning (d) Bi-tuning (b) Linear probing
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef4784a7-a11f-4472-a2d5-1cd6fa59181bCited by top-tier papers1
Ask how each one uses itBuilds on23
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- Co-Tuning for Transfer LearningKaichao You, Zhi Kou, Mingsheng Long, Jianmin WangNeurIPS 2020 · 105 citations
- Masked Images Are Counterfactual Samples for Robust Fine-TuningYao Xiao, Ziyi Tang, Pengxu Wei, Cong Liu et al.CVPR 2023
- Overwriting Pretrained Bias with Finetuning DataAngelina Wang, Olga RussakovskyICCV 2023 · 50 citations
- Preserving Commonsense Knowledge from Pre-trained Language Models via Causal InferenceJunhao Zheng, Qianli Ma, Shengjie Qiu, Yue Wu et al.ACL 2023 · 9 citations
- Causal Fine-Tuning under Latent Confounded ShiftJialin Yu, Yuxiang Zhou, Haoxuan Li, Junchi Yu et al.ICML 2026
