Robust fine-tuning of zero-shot models
Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, Ludwig Schmidt
Abstract
Large pre-trained models such as CLIP or ALIGN offer consistent accuracy across a range of data distributions when performing zero-shot inference (i.e., without fine-tuning on a specific dataset). Although existing fine-tuning methods substantially improve accuracy on a given target distribution, they often reduce robustness to distribution shifts. We address this tension by introducing a simple and effective method for improving robustness while fine-tuning: ensembling the weights of the zero-shot and fine-tuned models (WiSE-FT). Compared to standard fine-tuning, WiSE-FT provides large accuracy improvements under distribution shift, while preserving high accuracy on the target distribution. On ImageNet and five derived distribution shifts, WiSE-FT improves accuracy under distribution shift by 4 to 6 percentage points (pp) over prior work while increasing ImageNet accuracy by 1.6 pp. WiSE-FT achieves similarly large robustness gains (2 to 23 pp) on a diverse set of six further distribution shifts, and accuracy gains of 0.8 to 3.3 pp compared to standard fine-tuning on commonly used transfer learning datasets. These improvements come at no additional computational cost during fine-tuning or inference. Accuracy on the reference distribution (e.g., ImageNet) Accuracy on the distribution shifts M od el s tr ai ne d on re fe re nc e di st ri bu ti on tr ai n se t Z e ro -s h o t C L IP m o d e ls Effective robustness Fine-tuned CLIP Schematic: fine-tuning CLIP on the reference distribution leads to higher accuracy on the reference distribution but less robustness Accuracy on the reference distribution (e.g., ImageNet) Accuracy on the distribution shifts M od el s tr ai ne d on re fe re nc e di st ri bu ti on tr ai n se t Z e ro -s h o t C L IP m o d e ls Weight-space ensemble for α ∈ [0, 1]: θ α = (1α) • θ zero-shot + α • θ fine-tuned θ zero-shot θ fine-tuned Schematic: our method, WiSE-FT leads to better accuracy on the distribution shifts without decreasing accuracy on the reference distribution Var yingamixingac o e f f i c ie ntaα 55 60 65 70 75 80 85
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c340b3e7-61e5-4adc-b6d2-3d18b87769a8Cited by top-tier papers348
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma et al.ICLR 2022 · 911 citations
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
Related papers
- Robust Fine-tuning of Zero-shot Models via Variance ReductionBeier Zhu, Jiequan Cui, Hanwang ZhangNeurIPS 2024 · 9 citations
- Contrastive Adapters for Foundation Model Group RobustnessMichael Zhang, Christopher RéNeurIPS 2022 · 94 citations
- Lipsum-FT: Robust Fine-Tuning of Zero-Shot Models Using Random Text GuidanceGiung Nam, Byeongho Heo, Juho LeeICLR 2024 · 15 citations
- Models Out of Line: A Fourier Lens on Distribution Shift RobustnessSara Fridovich-Keil, Brian R. Bartoldson, James Diffenderfer, Bhavya Kailkhura et al.NeurIPS 2022 · 2 citations
- Learning Mask-aware CLIP Representations for Zero-Shot SegmentationSiyu Jiao, Yunchao Wei, Yaowei Wang, Yao Zhao et al.NeurIPS 2023 · 88 citations
