Learning Interpretable Queries for Explainable Image Classification with Information Pursuit
Stefan Kolek, Aditya Chattopadhyay, Kwan Ho Ryan Chan, Héctor Andrade-Loarca, Gitta Kutyniok, René Vidal
摘要
Information Pursuit (IP) is a recently introduced learning framework to construct classifiers that are interpretableby-design. Given a set of task-relevant and interpretable data queries, IP selects a small subset of the most informative queries and makes predictions based on the gathered query-answer pairs. However, a key limitation of IP is its dependency on task-relevant interpretable queries, which typically require considerable data annotation and curation efforts. While previous approaches have explored using general-purpose large language models to generate these query sets, they rely on prompt engineering heuristics and often yield suboptimal query sets, resulting in a performance gap between IP and non-interpretable blackbox predictors. In this work, we propose parameterizing IP queries as a learnable dictionary defined in the latent space of vision-language models such as CLIP. We formulate an optimization objective to learn IP queries and propose an alternating optimization algorithm that shares appealing connections with classic sparse dictionary learning algorithms. Our learned dictionary outperforms baseline methods based on handcrafted or prompted dictionaries across several image classification benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Conformal Information Pursuit for Interactively Guiding Large Language ModelsKwan Ho Ryan Chan, Yuyan Ge, Edgar Dobriban, Hamed Hassani 等NeurIPS 2025 · 被引用 9 次
- Testing Semantic Importance via BettingJacopo Teneggi, Jeremias SulamNeurIPS 2024 · 被引用 6 次
- Hierarchical Concept Embedding & Pursuit for Interpretable Image ClassificationNghia Nguyen, Tianjiao Ding, Rene VidalCVPR 2026 · 被引用 1 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Interpreting CLIP's Image Representation via Text-Based DecompositionYossi Gandelsman, Alexei A. Efros, Jacob SteinhardtICLR 2024 · 被引用 179 次
- Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)Usha Bhalla, Alex Oesterling, Suraj Srinivas, Flávio P. Calmon 等NeurIPS 2024 · 被引用 146 次
- Text-To-Concept (and Back) via Cross-Model AlignmentMazda Moayeri, Keivan Rezaei, Maziar Sanjabi, Soheil FeiziICML 2023 · 被引用 62 次
相关 Paper
- Bootstrapping Variational Information Pursuit with Large Language and Vision Models for Interpretable Image ClassificationAditya Chattopadhyay, Kwan Ho Ryan Chan, René VidalICLR 2024 · 被引用 12 次
- Variational Information Pursuit for Interpretable PredictionsAditya Chattopadhyay, Kwan Ho Ryan Chan, Benjamin David Haeffele, Donald Geman 等ICLR 2023
- IPO: Interpretable Prompt Optimization for Vision-Language ModelsYingjun Du, Wenfang Sun, Cees SnoekNeurIPS 2024 · 被引用 15 次
- What does CLIP know about a red circle? Visual prompt engineering for VLMsAleksandar Shtedritski, Christian Rupprecht, Andrea VedaldiICCV 2023 · 被引用 262 次
- Information Maximization Perspective of Orthogonal Matching Pursuit with Applications to Explainable AIAditya Chattopadhyay, Ryan Pilgrim, René VidalNeurIPS 2023 · 被引用 17 次
