Calibrating Multi-modal Representations: A Pursuit of Group Robustness without Annotations
Chenyu You, Yifei Min, Weicheng Dai, Jasjeet S. Sekhon, Lawrence H. Staib, James S. Duncan
摘要
Fine-tuning pre-trained vision-language models, like CLIP, has yielded success on diverse downstream tasks. However, several pain points persist for this paradigm: (i) directly tuning entire pre-trained models becomes both time-intensive and computationally costly. Additionally, these tuned models tend to become highly specialized, limiting their practicality for real-world deployment; (ii) recent studies indicate that pre-trained vision-language classifiers may overly depend on spurious features - patterns that correlate with the target in training data, but are not related to the true labeling function; and (iii) existing studies on mitigating the reliance on spurious features, largely based on the assumption that we can identify such features, does not provide definitive assurance for real-world applications. As a piloting study, this work focuses on exploring mitigating the reliance on spurious features for CLIP without using any group annotation. To this end, we systematically study the existence of spurious correlation on CLIP and CLIP+ERM. We first, following recent work on Deep Feature Reweighting (DFR), verify that last-layer retraining can greatly improve group robustness on pretrained CLIP. In view of them, we advocate a lightweight representation calibration method for fine-tuning CLIP, by first generating a calibration set using the pretrained CLIP, and then calibrating representations of samples within this set through contrastive learning, all without the need for group labels. Extensive experiments and in-depth visualizations on several benchmarks validate the effectiveness of our proposals, largely reducing reliance and significantly boosting the model generalization. Our codes will be available in here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- RSA: Resolving Scale Ambiguities in Monocular Depth Estimators through Language DescriptionsZiyao Zeng, Yangchao Wu, Hyoungseob Park, Daniel Wang 等NeurIPS 2024 · 被引用 26 次
- Why Prompt Design Matters and Works: A Complexity Analysis of Prompt Search Space in LLMsXiang Zhang, Juntai Cao, Chenyu You, Dujian DingACL 2025 · 被引用 21 次
- CSRv2: Unlocking Ultra-Sparse EmbeddingsLixuan Guo, Yifei Wang, Tiansheng Wen, Yifan Wang 等ICLR 2026 · 被引用 7 次
- HorizonForge: Driving Scene Editing with Any Trajectories and Any VehiclesYifan Wang, Francesco Pittaluga, Zaid Tasneem, Chenyu You 等CVPR 2026 · 被引用 3 次
- Decoupling Template Bias in CLIP: Harnessing Empty Prompts for Enhanced Few-Shot LearningZhenyu Zhang, Guangyao Chen, Yixiong Zou, Zhimeng Huang 等AAAI 2026 · 被引用 3 次
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
相关 Paper
- Mitigating Spurious Correlations in Multi-modal Models during Fine-tuningYu Yang, Besmira Nushi, Hamid Palangi, Baharan MirzasoleimanICML 2023 · 被引用 65 次
- Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation ModelHuan Ma, Yan Zhu, Changqing Zhang, Peilin Zhao 等AAAI 2025 · 被引用 4 次
- Density-Aware Translation of Spurious Correlations in Zero-Shot VLMsAfsaneh Hasanebrahimi, Hanxun Huang, Christopher Leckie, Sarah ErfaniICML 2026
- SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal BiasWenqian Ye, Di Wang, Guangtao Zheng, Bohan Liu 等AAAI 2026
- CPL: Counterfactual Prompt Learning for Vision and Language ModelsXuehai He, Diji Yang, Weixi Feng, Tsu-Jui Fu 等EMNLP 2022 · 被引用 13 次
