Conformal Prediction for Zero-Shot Models
Julio Silva-Rodríguez, Ismail Ben Ayed, Jose Dolz
摘要
Vision-language models pre-trained at large scale have shown unprecedented adaptability and generalization to downstream tasks. Although its discriminative potential has been widely explored, its reliability and uncertainty are still overlooked. In this work, we investigate the capabilities of CLIP models under the split conformal prediction paradigm, which provides theoretical guarantees to blackbox models based on a small, labeled calibration set. In contrast to the main body of literature on conformal predictors in vision classifiers, foundation models exhibit a particular characteristic: they are pre-trained on a one-time basis on an inaccessible source domain, different from the transferred task. This domain drift negatively affects the efficiency of the conformal sets and poses additional challenges. To alleviate this issue, we propose Conf-OT, a transfer learning setting that operates transductive over the combined calibration and query sets. Solving an optimal transport problem, the proposed method bridges the domain gap between pre-training and adaptation without requiring additional data splits but still maintaining coverage guarantees. We comprehensively explore this conformal prediction strategy on a broad span of 15 datasets and three nonconformity scores. Conf-OT provides consistent relative improvements of up to 20% on set efficiency while being ×15 faster than popular transductive approaches. We make the code available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation ModelsZhanxuan Hu, Qiyu Xu, Yu Duan, Yonghang Tai 等CVPR 2026 · 被引用 6 次
- LATA: Laplacian-Assisted Transductive Adaptation for Conformal Uncertainty in Medical VLMsBehzad Bozorgtabar, Dwarikanath Mahapatra, Sudipta Roy, Muzammal Naseer 等CVPR 2026 · 被引用 3 次
- Spectral Conformal Risk Control: Distribution-Free Tail Guarantees via Bayesian QuadratureMohammad Mahdi Kazemi Esfeh, Qi Yan, Yongxing Zhang, Zahra Gholami 等CVPR 2026
- CAOS: Conformal Aggregation of One-Shot PredictorsMaja WaldronICML 2026
它引用的顶会 Paper35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 被引用 873 次
相关 Paper
- Leveraging Vision-Language Models for Improving Domain Generalization in Image ClassificationSravanti Addepalli, Ashish Ramayee Asokan, Lakshay Sharma, R. Venkatesh BabuCVPR 2024
- TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent CollaborationYiwei Guo, Shaobin Zhuang, Kunchang Li, Yu Qiao 等NeurIPS 2024 · 被引用 9 次
- Inverse Optimal Transport for Efficient Adaptation of Vision-Language ModelsShupeng Qiu, Chuan-Xian RenAAAI 2026 · 被引用 1 次
- Towards Calibrating Prompt Tuning of Vision- Language ModelsAshshak Sharifdeen, Fahad Shamshad, Muhammad Akhtar Munir, Abhishek Basu 等CVPR 2026
- Boosting Vision-Language Models with TransductionMaxime Zanella, Benoît Gérin, Ismail Ben AyedNeurIPS 2024 · 被引用 42 次
