Towards Calibrated Robust Fine-Tuning of Vision-Language Models
Changdae Oh, Hyesu Lim, Mijoo Kim, Dongyoon Han, Sangdoo Yun, Jaegul Choo, Alexander Hauptmann, Zhi-Qi Cheng, Kyungwoo Song
摘要
Improving out-of-distribution (OOD) generalization during in-distribution (ID) adaptation is a primary goal of robust fine-tuning of zero-shot models beyond naive fine-tuning. However, despite decent OOD generalization performance from recent robust fine-tuning methods, confidence calibration for reliable model output has not been fully addressed. This work proposes a robust fine-tuning method that improves both OOD accuracy and confidence calibration simultaneously in vision language models. Firstly, we show that both OOD classification and OOD calibration errors have a shared upper bound consisting of two terms of ID data: 1) ID calibration error and 2) the smallest singular value of the ID input covariance matrix. Based on this insight, we design a novel framework that conducts fine-tuning with a constrained multimodal contrastive loss enforcing a larger smallest singular value, which is further guided by the self-distillation of a moving-averaged model to achieve calibrated prediction as well. Starting from empirical evidence supporting our theoretical statements, we provide extensive experimental results on ImageNet distribution shift benchmarks that demonstrate the effectiveness of our theorem and its practical implementation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution DetectionGeng Yu, Jianing Zhu, Jiangchao Yao, Bo HanNeurIPS 2024 · 被引用 26 次
- Open-Vocabulary Calibration for Fine-tuned CLIPShuoyuan Wang, Jindong Wang, Guoqing Wang, Bob Zhang 等ICML 2024 · 被引用 17 次
- Understanding Language Prior of LVLMs by Contrasting Chain-of-EmbeddingLin Long, Changdae Oh, Seongheon Park, Sharon LiICLR 2026 · 被引用 14 次
- MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge TransferMinghao Zhu, Zhengpu Wang, Mengxian Hu, Ronghao Dang 等NeurIPS 2024 · 被引用 10 次
- Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and OpportunitiesChangdae Oh, Seongheon Park, To Eun Kim, Jiatong Li 等ACL 2026 · 被引用 8 次
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu 等ICML 2022 · 被引用 1,123 次
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma 等ICLR 2022 · 被引用 911 次
相关 Paper
- AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language ModelsYubo Cui, Xianchao Guan, Zijun Xiong, Zheng ZhangCVPR 2026 · 被引用 1 次
- Lipsum-FT: Robust Fine-Tuning of Zero-Shot Models Using Random Text GuidanceGiung Nam, Byeongho Heo, Juho LeeICLR 2024 · 被引用 15 次
- Difference Vector Equalization for Robust Fine-tuning of Vision-Language ModelsSatoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane 等AAAI 2026
- TRACER: Persistent Regularization for Robust Multimodal FinetuningHesam Asadollahzadeh, Feng Liu, Christopher Leckie, Sarah ErfaniICML 2026
- Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD GeneralizationYuhang Zang, Hanlin Goh, Joshua Susskind, Chen HuangICLR 2024 · 被引用 17 次
