A Closer Look at the Few-Shot Adaptation of Large Vision-Language Models
Julio Silva-Rodríguez, Sina Hajimiri, Ismail Ben Ayed, Jose Dolz
摘要
Efficient transfer learning (ETL) is receiving increasing attention to adapt large pre-trained language-vision models on downstream tasks with a few labeled samples. While significant progress has been made, we reveal that state-of-the-art ETL approaches exhibit strong performance only in narrowly-defined experimental setups, and with a careful adjustment of hyperparameters based on a large corpus of labeled samples. In particular, we make two interesting, and surprising empirical observations. First, to out-perform a simple Linear Probing baseline, these methods require to optimize their hyper-parameters on each target task. And second, they typically underperform -sometimes dramatically- standard zero-shot predictions in the presence of distributional drifts. Motivated by the unrealistic assumptions made in the existing literature, i.e., access to a large validation set and case-specific grid-search for optimal hyperparameters, we propose a novel approach that meets the requirements of real-world scenarios. More concretely, we introduce a CLass-Adaptive linear Probe (CLAP) objective, whose balancing term is optimized via an adaptation of the general Augmented Lagrangian method tailored to this context. We comprehensively evaluate CLAP on a broad span of datasets and scenarios, demonstrating that it consistently outperforms SoTA approaches, while yet being a much more efficient alternative. Code available at https://github.com/jusiro/CLAP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Text and Image Are Mutually Beneficial: Enhancing Training-Free Few-Shot Classification with CLIPYayuan Li, Jintao Guo, Lei Qi, Wenbin Li 等AAAI 2025 · 被引用 9 次
- FLOSS: Free Lunch in Open-Vocabulary Semantic SegmentationYasser Benigmim, Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc 等ICCV 2025 · 被引用 6 次
- LATA: Laplacian-Assisted Transductive Adaptation for Conformal Uncertainty in Medical VLMsBehzad Bozorgtabar, Dwarikanath Mahapatra, Sudipta Roy, Muzammal Naseer 等CVPR 2026 · 被引用 3 次
- Dual-Kernel Adapter: Expanding Spatial Horizons for Data-Constrained Medical Image AnalysisZiquan Zhu, Hanruo Zhu, Si-Yuan Lu, Xiang Li 等ICLR 2026 · 被引用 3 次
- Towards Effective Foundation Model Adaptation for Extreme Cross-Domain Few-Shot LearningFei Zhou, Peng Wang, Lei Zhang, Wei Wei 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
相关 Paper
- Towards Difficulty-Agnostic Efficient Transfer Learning for Vision-Language ModelsYongjin Yang, Jongwoo Ko, Se-Young YunEMNLP 2024 · 被引用 1 次
- Project and Probe: Sample-Efficient Adaptation by Interpolating Orthogonal FeaturesAnnie S. Chen, Yoonho Lee, Amrith Setlur, Sergey Levine 等ICLR 2024 · 被引用 5 次
- Rethinking the Effect of Uninformative Class Name in Prompt LearningFengmao Lv, Changru Nie, Jianyang Zhang, Guowu Yang 等ACM MM 2024 · 被引用 1 次
- Efficient and Long-Tailed Generalization for Pre-trained Vision-Language ModelJiang-Xin Shi, Chi Zhang, Tong Wei, Yufeng LiKDD 2024 · 被引用 3 次
- Task Residual for Tuning Vision-Language ModelsTao Yu, Zhihe Lu, Xin Jin, Zhibo Chen 等CVPR 2023
