Interpretable Image Classification with Adaptive Prototype-based Vision Transformers
Chiyu Ma, Jon Donnelly, Wenjun Liu, Soroush Vosoughi, Cynthia Rudin, Chaofan Chen
Abstract
We present ProtoViT, a method for interpretable image classification combining deep learning and case-based reasoning. This method classifies an image by comparing it to a set of learned prototypes, providing explanations of the form "this looks like that." In our model, a prototype consists of parts, which can deform over irregular geometries to create a better comparison between images. Unlike existing models that rely on Convolutional Neural Network (CNN) backbones and spatially rigid prototypes, our model integrates Vision Transformer (ViT) backbones into prototype based models, while offering spatially deformed prototypes that not only accommodate geometric variations of objects but also provide coherent and clear prototypical feature representations with an adaptive number of prototypical parts. Our experiments show that our model can generally achieve higher performance than the existing prototype based models. Our comprehensive analyses ensure that the prototypes are consistent and the interpretations are faithful. Our code is available at https://github.com/Henrymachiyu/ProtoViT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 160de91e-615d-450d-be13-e04cbf2970f3Cited by top-tier papers9
- ProtoLens: Advancing Prototype Learning for Fine-Grained Interpretability in Text ClassificationBowen Wei, Ziwei ZhuACL 2025 · 7 citations
- Debugging Concept Bottleneck Models through Removal and RetrainingEric Enouen, Sainyam GalhotraICLR 2026 · 2 citations
- Prototype Transformer: Towards Language Model Architectures Interpretable by DesignYordan Yordanov, Matteo Forasassi, Bayar Menzat, Ruizhi Wang et al.ICML 2026 · 1 citation
- DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing ModalitiesJueqing Lu, Yuanyuan Qi, Xiaohao Yang, Shuaicheng Niu et al.CVPR 2026
- From Pixels to Perception: Interpretable Predictions via Instance-wise Grouped Feature SelectionMoritz Vandenhirtz, Julia E. VogtICML 2025
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
Related papers
- Deformable ProtoPNet: An Interpretable Image Classifier Using Deformable PrototypesJon Donnelly, Alina Jade Barnett, Chaofan ChenCVPR 2022 · 101 citations
- This Looks Like Those: Illuminating Prototypical Concepts Using Multiple VisualizationsChiyu Ma, Brandon Zhao, Chaofan Chen, Cynthia RudinNeurIPS 2023 · 53 citations
- ProtoArgNet: Interpretable Image Classification with Super-Prototypes and ArgumentationHamed Ayoobi, Nico Potyka, Francesca ToniAAAI 2025 · 8 citations
- This Looks Like It Rather Than That: ProtoKNN For Similarity-Based ClassifiersYuki Ukai, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu FujiyoshiICLR 2023
- Interpretable Image Classification via Non-parametric Part Prototype LearningZhijie Zhu, Lei Fan, Maurice Pagnucco, Yang SongCVPR 2025
