Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
Shiming Chen, Bowen Duan, Salman Khan, Fahad Shahbaz Khan
摘要
Large-scale vision-language models (VLMs), such as CLIP, have achieved remarkable success in zero-shot learning (ZSL) by leveraging large-scale visual-text pair datasets. However, these methods often lack interpretability, as they compute the similarity between an entire query image and the embedded category words, making it difficult to explain their predictions. One approach to address this issue is to develop interpretable models by integrating language, where classifiers are built using discrete attributes, similar to human perception. This introduces a new challenge: how to effectively align local visual features with corresponding attributes based on pre-trained VLMs. To tackle this, we propose LaZSL, a locally-aligned vision-language model for interpretable ZSL. LaZSL employs local visual-semantic alignment via optimal transport to perform interaction between visual regions and their associated attributes, facilitating effective alignment and providing interpretable similarity without the need for additional training. Extensive experiments demonstrate that our method offers several advantages, including enhanced interpretability, improved accuracy, and strong domain generalization. Codes available at: https://github.com/shiming-chen/ LaZSL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LOREAL: Mitigating Low-Resolution Challenges in Vision-Language Models with Attribute-driven Prompt Self-DistillationXucong Wang, Pengkun Wang, Zhe Zhao, Liheng Yu 等CVPR 2026
- Learning to Memorize with Attributive and Associative Memory for Online Test-Time Adaptation of Vision-Language ModelsYuchao Zhang, Hao Wang, Fan Zhang, QIRUI MI 等ICML 2026
- FedMPT: Federated Multi-Label Prompt Tuning of Vision-Language ModelsXucong Wang, Pengkun Wang, Zhe Zhao, Liheng Yu 等CVPR 2026
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu 等NeurIPS 2022 · 被引用 603 次
相关 Paper
- Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language ModelsJinhao Li, Haopeng Li, Sarah Monazam Erfani, Lei Feng 等ICML 2024 · 被引用 30 次
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang 等ICML 2024 · 被引用 7 次
- From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based SelectionLincan Cai, Jingxuan Kang, Shuang Li, Wenxuan Ma 等ICML 2025
- Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic SegmentationYunheng Li, Zhong-Yu Li, Quan-Sheng Zeng, Qibin Hou 等ICML 2024 · 被引用 27 次
- Unbiased Region-Language Alignment for Open-Vocabulary Dense PredictionYunheng Li, Yuxuan Li, Quan-Sheng Zeng, Wenhai Wang 等ICCV 2025 · 被引用 3 次
