Bootstrapping Variational Information Pursuit with Large Language and Vision Models for Interpretable Image Classification
Aditya Chattopadhyay, Kwan Ho Ryan Chan, René Vidal
摘要
Variational Information Pursuit (V-IP) is an interpretable-by-design framework that makes predictions by sequentially selecting a short chain of user-defined, interpretable queries about the data that are most informative for the task. The prediction is based solely on the obtained query answers, which also serve as a faithful explanation for the prediction. Applying the framework to any task requires (i) specification of a query set, and (ii) densely annotated data with query answers to train classifiers to answer queries at test time. This limits V-IP's application to small-scale tasks where manual data annotation is feasible. In this work, we focus on image classification tasks and propose to relieve this bottleneck by leveraging pretrained language and vision models. Specifically, following recent work, we propose to use GPT, a Large Language Model, to propose semantic concepts as queries for a given classification task. To answer these queries, we propose a light-weight Concept Question-Answering network (Concept-QA) which learns to answer binary queries about semantic concepts in images. We design pseudo-labels to train our Concept-QA model using GPT and CLIP (a Vision-Language Model). Empirically, we find our Concept-QA model to be competitive with state-of-the-art VQA models in terms of answering accuracy but with an order of magnitude fewer parameters. This allows for seamless integration of Concept-QA into the V-IP framework as a fast-answering mechanism. We name this method Concept-QA+V-IP. Finally, we show on several datasets that Concept-QA+V-IP produces shorter, interpretable query chains which are more accurate than V-IP trained with CLIP-based answering systems. Code available at https://github.com/adityac94/conceptqa_vip .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation ModelsHengyi Wang, Shiwei Tan, Hao WangICML 2024 · 被引用 9 次
- Conformal Information Pursuit for Interactively Guiding Large Language ModelsKwan Ho Ryan Chan, Yuyan Ge, Edgar Dobriban, Hamed Hassani 等NeurIPS 2025 · 被引用 9 次
- Testing Semantic Importance via BettingJacopo Teneggi, Jeremias SulamNeurIPS 2024 · 被引用 6 次
- Hierarchical Concept Embedding & Pursuit for Interpretable Image ClassificationNghia Nguyen, Tianjiao Ding, Rene VidalCVPR 2026 · 被引用 1 次
- ImageSet2Text: Describing Sets of Images Through TextPiera Riccio, Francesco Galati, Kajetan Schweighofer, Noa Garcia 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Learning Interpretable Queries for Explainable Image Classification with Information PursuitStefan Kolek, Aditya Chattopadhyay, Kwan Ho Ryan Chan, Héctor Andrade-Loarca 等ICCV 2025 · 被引用 1 次
- Variational Information Pursuit for Interpretable PredictionsAditya Chattopadhyay, Kwan Ho Ryan Chan, Benjamin David Haeffele, Donald Geman 等ICLR 2023
- Align2Concept: Language Guided Interpretable Image Recognition by Visual Prototype and Textual Concept AlignmentJiaqi Wang, Pichao Wang, Yi Feng, Huafeng Liu 等ACM MM 2024 · 被引用 1 次
- V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept TokenizerHangzhou He, Lei Zhu, Xinliang Zhang, Shuang Zeng 等AAAI 2025 · 被引用 11 次
- When are Lemons Purple? The Concept Association Bias of Vision-Language ModelsYingtian Tang, Yutaro Yamada, Yoyo Zhang, Ilker YildirimEMNLP 2023 · 被引用 9 次
