Leveraging Data to Say No: Memory Augmented Plug-and-Play Selective Prediction
Aditya Sarkar, Yi Li, Jiacheng Cheng, Shlok Kumar Mishra, Nuno Vasconcelos
Abstract
Selective prediction aims to endow predictors with a reject option, to avoid low confidence predictions. However, existing literature has primarily focused on closed-set tasks, such as visual question answering with predefined options or fixed-category classification. This paper considers selective prediction for visual language foundation models, addressing a taxonomy of tasks ranging from closed to open set and from finite to unbounded vocabularies, as in image captioning. We seek training-free approaches of low-complexity, applicable to any foundation model and consider methods based on external vision-language model (VLM) embeddings, like CLIP. This is denoted as \textit{Plug-and-Play Selective Prediction} (\textbf{\texttt{PaPSP}}). We identify two key challenges: (1) , leading to high variance in image-text embeddings, and (2) . To address these issues, we propose a \textbf{\texttt{PaPSP}} (\textbf{\texttt{MA-PaPSP}}) model, which augments \textbf{\texttt{PaPSP}} with a retrieval dataset of image-text pairs. This is leveraged to reduce embedding variance by averaging retrieved nearest-neighbor pairs and is complemented by the use of contrastive normalization to improve score calibration. Through extensive experiments on multiple datasets, we show that \textbf{\texttt{MA-PaPSP}} outperforms \textbf{\texttt{PaPSP}} and other selective prediction baselines for selective captioning, image-text matching, and fine-grained classification. Source code will be made public.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
Related papers
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-languageDelong Chen, Mustafa Shukor, Théo Moutakanni, Willy Chung et al.ICLR 2026 · 60 citations
- From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based SelectionLincan Cai, Jingxuan Kang, Shuang Li, Wenxuan Ma et al.ICML 2025
- CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic SegmentationSeokju Cho, Heeseong Shin, Sunghwan Hong, Anurag Arnab et al.CVPR 2024
- DeCLIP: Decoupled Learning for Open-Vocabulary Dense PerceptionJunjie Wang, Bin Chen, Yulin Li, Bin Kang et al.CVPR 2025
- MaskInversion: Localized Embeddings via Optimization of Explainability MapsWalid Bousselham, Sofian Chaybouti, Christian Rupprecht, Vittorio Ferrari et al.ICLR 2026 · 3 citations
