Bootstrapping Variational Information Pursuit with Large Language and Vision Models for Interpretable Image Classification
Aditya Chattopadhyay, Kwan Ho Ryan Chan, René Vidal
Abstract
Variational Information Pursuit (V-IP) is an interpretable-by-design framework that makes predictions by sequentially selecting a short chain of user-defined, interpretable queries about the data that are most informative for the task. The prediction is based solely on the obtained query answers, which also serve as a faithful explanation for the prediction. Applying the framework to any task requires (i) specification of a query set, and (ii) densely annotated data with query answers to train classifiers to answer queries at test time. This limits V-IP's application to small-scale tasks where manual data annotation is feasible. In this work, we focus on image classification tasks and propose to relieve this bottleneck by leveraging pretrained language and vision models. Specifically, following recent work, we propose to use GPT, a Large Language Model, to propose semantic concepts as queries for a given classification task. To answer these queries, we propose a light-weight Concept Question-Answering network (Concept-QA) which learns to answer binary queries about semantic concepts in images. We design pseudo-labels to train our Concept-QA model using GPT and CLIP (a Vision-Language Model). Empirically, we find our Concept-QA model to be competitive with state-of-the-art VQA models in terms of answering accuracy but with an order of magnitude fewer parameters. This allows for seamless integration of Concept-QA into the V-IP framework as a fast-answering mechanism. We name this method Concept-QA+V-IP. Finally, we show on several datasets that Concept-QA+V-IP produces shorter, interpretable query chains which are more accurate than V-IP trained with CLIP-based answering systems. Code available at https://github.com/adityac94/conceptqa_vip .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95d08b41-ac99-41aa-b936-b625896c8618Cited by top-tier papers7
- Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation ModelsHengyi Wang, Shiwei Tan, Hao WangICML 2024 · 9 citations
- Conformal Information Pursuit for Interactively Guiding Large Language ModelsKwan Ho Ryan Chan, Yuyan Ge, Edgar Dobriban, Hamed Hassani et al.NeurIPS 2025 · 9 citations
- Testing Semantic Importance via BettingJacopo Teneggi, Jeremias SulamNeurIPS 2024 · 6 citations
- Hierarchical Concept Embedding & Pursuit for Interpretable Image ClassificationNghia Nguyen, Tianjiao Ding, Rene VidalCVPR 2026 · 1 citation
- ImageSet2Text: Describing Sets of Images Through TextPiera Riccio, Francesco Galati, Kajetan Schweighofer, Noa Garcia et al.AAAI 2026 · 1 citation
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Learning Interpretable Queries for Explainable Image Classification with Information PursuitStefan Kolek, Aditya Chattopadhyay, Kwan Ho Ryan Chan, Héctor Andrade-Loarca et al.ICCV 2025 · 1 citation
- Variational Information Pursuit for Interpretable PredictionsAditya Chattopadhyay, Kwan Ho Ryan Chan, Benjamin David Haeffele, Donald Geman et al.ICLR 2023
- Align2Concept: Language Guided Interpretable Image Recognition by Visual Prototype and Textual Concept AlignmentJiaqi Wang, Pichao Wang, Yi Feng, Huafeng Liu et al.ACM MM 2024 · 1 citation
- V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept TokenizerHangzhou He, Lei Zhu, Xinliang Zhang, Shuang Zeng et al.AAAI 2025 · 11 citations
- When are Lemons Purple? The Concept Association Bias of Vision-Language ModelsYingtian Tang, Yutaro Yamada, Yoyo Zhang, Ilker YildirimEMNLP 2023 · 9 citations
