SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
Kevin Miller, Aditya Gangrade, Samarth Mishra, Kate Saenko, Venkatesh Saligrama
摘要
Zero-shot multi-label recognition (MLR) with Vision-Language Models (VLMs) faces significant challenges without training data, model tuning, or architectural modifications. Existing approaches require prompt tuning or architectural adaptations, limiting zero-shot applicability. Our work proposes a novel solution treating VLMs as black boxes, leveraging scores without training data or ground truth. We make two contributions. First, we find that VLM scores suffer from image-and prompt-specific biases, and that simple standardization is surprisingly effective at removing these and boosting MLR performance. And second, we introduce compound prompts grounded in realistic object combinations. Our analysis reveals "AND"/"OR" signal ambiguities that cause maximum compound scores to be surprisingly suboptimal compared to second-highest scores. We introduce an adaptive fusion method to address this issue. Our method enhances other zero-shot approaches, consistently improving their results. Experiments show superior mean Average Precision (mAP) compared to methods requiring training data, achieved through refined object ranking for robust zero-shot MLR. Code can be found at https://github.com/kjmillerCURIS/SPARC .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- [CLS] is Not Enough: Multi-Label Recognition via Patch-Level Inference and Adaptive AggregationAkang Wang, Xili Deng, Zhanxuan Hu, Yi Zhao 等ICML 2026 · 被引用 1 次
- Multi-Label Test-Time Adaptation with Bayesian Conditional PriorsQiru Li, Ao Zhou, Zhiwei Jiang, Zifeng Cheng 等ICML 2026 · 被引用 1 次
- Rethinking BCE Loss for Multi-Label Image Recognition with Fine-TuningAo Zhou, Zhiwei Jiang, Zifeng Cheng, Cong Wang 等CVPR 2026
- FedMPT: Federated Multi-Label Prompt Tuning of Vision-Language ModelsXucong Wang, Pengkun Wang, Zhe Zhao, Liheng Yu 等CVPR 2026
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- What does a platypus look like? Generating customized prompts for zero-shot image classificationSarah M. Pratt, Ian Covert, Rosanne Liu, Ali FarhadiICCV 2023 · 被引用 343 次
- DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited AnnotationsXimeng Sun, Ping Hu, Kate SaenkoNeurIPS 2022 · 被引用 199 次
- CHiLS: Zero-Shot Image Classification with Hierarchical Label SetsZachary Novack, Julian J. McAuley, Zachary Chase Lipton, Saurabh GargICML 2023 · 被引用 127 次
- CDUL: CLIP-Driven Unsupervised Learning for Multi-Label Image ClassificationRabab Abdelfattah, Qing Guo, Xiaoguang Li, Xiaofeng Wang 等ICCV 2023 · 被引用 58 次
相关 Paper
- CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Vision-Language ModelRuijiang Dong, Zesheng Ye, Jianzhong Qi, Lei Feng 等ICLR 2026
- Mitigating Spurious Correlations in Zero-Shot Multimodal ModelsShenyu Lu, Junyi Chai, Xiaoqian WangICLR 2025
- Decomposed Soft Prompt Guided Fusion Enhancing for Compositional Zero-Shot LearningXiaocheng Lu, Song Guo, Ziming Liu, Jingcai GuoCVPR 2023
- Revisiting the Role of Language Priors in Vision-Language ModelsZhiqiu Lin, Xinyue Chen, Deepak Pathak, Pengchuan Zhang 等ICML 2024 · 被引用 44 次
- LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Models for Referring Expression ComprehensionAmaia Cardiel, Eloi Zablocki, Elias Ramzi, Oriane Siméoni 等ICLR 2025
