Improving Zero-shot Generalization and Robustness of Multi-Modal Models
Yunhao Ge, Jie Ren, Andrew Gallagher, Yuxiao Wang, Ming-Hsuan Yang, Hartwig Adam, Laurent Itti, Balaji Lakshminarayanan, Jiaping Zhao
摘要
Multi-modal image-text models such as CLIP and LiT have demonstrated impressive performance on image classification benchmarks and their zero-shot generalization ability is particularly exciting. While the top-5 zero-shot accuracies of these models are very high, the top-1 accuracies are much lower (over 25% gap in some cases). We investigate the reasons for this performance gap and find that many of the failure cases are caused by ambiguity in the text prompts. First, we develop a simple and efficient zero-shot post-hoc method to identify images whose top-1 prediction is likely to be incorrect, by measuring consistency of the predictions w.r.t. multiple prompts and image transformations. We show that our procedure better predicts mistakes, outperforming the popular max logit baseline on selective prediction tasks. Next, we propose a simple and efficient way to improve accuracy on such uncertain images by making use of the WordNet hierarchy; specifically we augment the original class by incorporating its parent and children from the semantic label hierarchy, and plug the augmentation into text prompts. We conduct experiments on both CLIP and LiT models with five different ImageNetbased datasets. For CLIP, our method improves the top-1 accuracy by 17.13% on the uncertain subset and 3.6% on the entire ImageNet validation set. We also show that our method improves across ImageNet shifted datasets, four other datasets, and other model architectures such as LiT. The proposed method 1 is hyperparameter-free, requires no additional model training and can be easily scaled to other large multi-modal architectures. Code is available at https://github.com/gyhandy/Hierarchy-CLIP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- TinyCLIP: CLIP Distillation via Affinity Mimicking and Weight InheritanceKan Wu, Houwen Peng, Zhenghong Zhou, Bin Xiao 等ICCV 2023 · 被引用 118 次
- A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image ModelsJames Urquhart Allingham, Jie Ren, Michael W. Dusenberry, Xiuye Gu 等ICML 2023 · 被引用 63 次
- AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationYuhan Zhu, Yuyang Ji, Zhiyu Zhao, Gangshan Wu 等NeurIPS 2024 · 被引用 45 次
- Bridge the Modality and Capability Gaps in Vision-Language Model SelectionChao Yi, Yuhang He, De-Chuan Zhan, Han-Jia YeNeurIPS 2024 · 被引用 32 次
- Democratizing Fine-grained Visual Recognition with Large Language ModelsMingxuan Liu, Subhankar Roy, Wenjing Li, Zhun Zhong 等ICLR 2024 · 被引用 27 次
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
相关 Paper
- Language-Driven Multi-Label Zero-Shot Learning with Semantic GranularityShouwen Wang, Qian Wan, Junbin Gao, Zhigang ZengICCV 2025 · 被引用 2 次
- CHiLS: Zero-Shot Image Classification with Hierarchical Label SetsZachary Novack, Julian J. McAuley, Zachary Chase Lipton, Saurabh GargICML 2023 · 被引用 127 次
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu 等NeurIPS 2022 · 被引用 603 次
- ChatGPT-Powered Hierarchical Comparisons for Image ClassificationZhiyuan Ren, Yiyang Su, Xiaoming LiuNeurIPS 2023 · 被引用 54 次
- SoftCLIP: Softer Cross-Modal Alignment Makes CLIP StrongerYuting Gao, Jinfeng Liu, Zihan Xu, Tong Wu 等AAAI 2024 · 被引用 80 次
