PRISM: Prototype-based Reasoning with Inter-modal Semantic Mining for Interpretable Image Recognition
Anni Yu, Yu-Bin Yang
Abstract
Prototype-based methods enhance interpretability in image recognition by establishing intermediate part prototypes to build interpretable classifiers, enabling transparent reasoning through part-level attention and reference to prototypical examples. However, existing methods typically depend on unimodal visual supervision and constrain prototypes within the visual embedding space, which inherently restricts their semantic alignment with human-interpretable concepts. In this work, we present PRISM (Prototype-based Reasoning with Inter-modal Semantic Mining), an interpretable image recognition framework that leverages natural language as an auxiliary modality to guide the learning of class-specific part prototypes. PRISM introduces an information-theoretic attribution mechanism that identifies semantically salient image regions conditioned on textual descriptions. By aligning these attribution maps with prototype activation patterns, PRISM implicitly anchors visual part prototypes to conceptually meaningful image regions, enhancing interpretability without requiring explicit concept modeling. To further enhance the distinctiveness and localization of prototypes, we introduce a spatial compactness constraint that encourages each prototype to attend to specific, non-overlapping image regions. Extensive experiments on fine-grained benchmarks demonstrate that the proposed PRISM not only improves classification performance but also provides faithful and semantically grounded visual explanations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65553615-48e3-4e77-9670-7d59cddf3f41Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Generating Images with Multimodal Language ModelsJing Yu Koh, Daniel Fried, Russ SalakhutdinovNeurIPS 2023 · 403 citations
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 220 citations
- Interpretable Image Recognition by Constructing Transparent Embedding SpaceJiaqi Wang, Huafeng Liu, Xinyue Wang, Liping JingICCV 2021 · 149 citations
Related papers
- Prototype-Guided Multimodal Relation Extraction based on Entity AttributesZefan Zhang, Weiqi Zhang, Yanhui Li, Tian BaiAAAI 2025 · 8 citations
- PIP-Net: Patch-Based Intuitive Prototypes for Interpretable Image ClassificationMeike Nauta, Jörg Schlötterer, Maurice van Keulen, Christin SeifertCVPR 2023
- Align2Concept: Language Guided Interpretable Image Recognition by Visual Prototype and Textual Concept AlignmentJiaqi Wang, Pichao Wang, Yi Feng, Huafeng Liu et al.ACM MM 2024 · 1 citation
- Interpretable Image Classification via Non-parametric Part Prototype LearningZhijie Zhu, Lei Fan, Maurice Pagnucco, Yang SongCVPR 2025
- Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual KnowledgeFawaz Sammani, Nikos DeligiannisNeurIPS 2024 · 15 citations
