Learning Interpretable Queries for Explainable Image Classification with Information Pursuit
Stefan Kolek, Aditya Chattopadhyay, Kwan Ho Ryan Chan, Héctor Andrade-Loarca, Gitta Kutyniok, René Vidal
Abstract
Information Pursuit (IP) is a recently introduced learning framework to construct classifiers that are interpretableby-design. Given a set of task-relevant and interpretable data queries, IP selects a small subset of the most informative queries and makes predictions based on the gathered query-answer pairs. However, a key limitation of IP is its dependency on task-relevant interpretable queries, which typically require considerable data annotation and curation efforts. While previous approaches have explored using general-purpose large language models to generate these query sets, they rely on prompt engineering heuristics and often yield suboptimal query sets, resulting in a performance gap between IP and non-interpretable blackbox predictors. In this work, we propose parameterizing IP queries as a learnable dictionary defined in the latent space of vision-language models such as CLIP. We formulate an optimization objective to learn IP queries and propose an alternating optimization algorithm that shares appealing connections with classic sparse dictionary learning algorithms. Our learned dictionary outperforms baseline methods based on handcrafted or prompted dictionaries across several image classification benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Conformal Information Pursuit for Interactively Guiding Large Language ModelsKwan Ho Ryan Chan, Yuyan Ge, Edgar Dobriban, Hamed Hassani et al.NeurIPS 2025 · 9 citations
- Testing Semantic Importance via BettingJacopo Teneggi, Jeremias SulamNeurIPS 2024 · 6 citations
- Hierarchical Concept Embedding & Pursuit for Interpretable Image ClassificationNghia Nguyen, Tianjiao Ding, Rene VidalCVPR 2026 · 1 citation
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Interpreting CLIP's Image Representation via Text-Based DecompositionYossi Gandelsman, Alexei A. Efros, Jacob SteinhardtICLR 2024 · 179 citations
- Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)Usha Bhalla, Alex Oesterling, Suraj Srinivas, Flávio P. Calmon et al.NeurIPS 2024 · 146 citations
- Text-To-Concept (and Back) via Cross-Model AlignmentMazda Moayeri, Keivan Rezaei, Maziar Sanjabi, Soheil FeiziICML 2023 · 62 citations
Related papers
- Bootstrapping Variational Information Pursuit with Large Language and Vision Models for Interpretable Image ClassificationAditya Chattopadhyay, Kwan Ho Ryan Chan, René VidalICLR 2024 · 12 citations
- Variational Information Pursuit for Interpretable PredictionsAditya Chattopadhyay, Kwan Ho Ryan Chan, Benjamin David Haeffele, Donald Geman et al.ICLR 2023
- IPO: Interpretable Prompt Optimization for Vision-Language ModelsYingjun Du, Wenfang Sun, Cees SnoekNeurIPS 2024 · 15 citations
- What does CLIP know about a red circle? Visual prompt engineering for VLMsAleksandar Shtedritski, Christian Rupprecht, Andrea VedaldiICCV 2023 · 262 citations
- Information Maximization Perspective of Orthogonal Matching Pursuit with Applications to Explainable AIAditya Chattopadhyay, Ryan Pilgrim, René VidalNeurIPS 2023 · 17 citations
