Interpretable Image Classification with Adaptive Prototype-based Vision Transformers
Chiyu Ma, Jon Donnelly, Wenjun Liu, Soroush Vosoughi, Cynthia Rudin, Chaofan Chen
摘要
We present ProtoViT, a method for interpretable image classification combining deep learning and case-based reasoning. This method classifies an image by comparing it to a set of learned prototypes, providing explanations of the form "this looks like that." In our model, a prototype consists of parts, which can deform over irregular geometries to create a better comparison between images. Unlike existing models that rely on Convolutional Neural Network (CNN) backbones and spatially rigid prototypes, our model integrates Vision Transformer (ViT) backbones into prototype based models, while offering spatially deformed prototypes that not only accommodate geometric variations of objects but also provide coherent and clear prototypical feature representations with an adaptive number of prototypical parts. Our experiments show that our model can generally achieve higher performance than the existing prototype based models. Our comprehensive analyses ensure that the prototypes are consistent and the interpretations are faithful. Our code is available at https://github.com/Henrymachiyu/ProtoViT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- ProtoLens: Advancing Prototype Learning for Fine-Grained Interpretability in Text ClassificationBowen Wei, Ziwei ZhuACL 2025 · 被引用 7 次
- Debugging Concept Bottleneck Models through Removal and RetrainingEric Enouen, Sainyam GalhotraICLR 2026 · 被引用 2 次
- Prototype Transformer: Towards Language Model Architectures Interpretable by DesignYordan Yordanov, Matteo Forasassi, Bayar Menzat, Ruizhi Wang 等ICML 2026 · 被引用 1 次
- DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing ModalitiesJueqing Lu, Yuanyuan Qi, Xiaohao Yang, Shuaicheng Niu 等CVPR 2026
- From Pixels to Perception: Interpretable Predictions via Instance-wise Grouped Feature SelectionMoritz Vandenhirtz, Julia E. VogtICML 2025
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,279 次
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 被引用 480 次
相关 Paper
- Deformable ProtoPNet: An Interpretable Image Classifier Using Deformable PrototypesJon Donnelly, Alina Jade Barnett, Chaofan ChenCVPR 2022 · 被引用 101 次
- This Looks Like Those: Illuminating Prototypical Concepts Using Multiple VisualizationsChiyu Ma, Brandon Zhao, Chaofan Chen, Cynthia RudinNeurIPS 2023 · 被引用 53 次
- ProtoArgNet: Interpretable Image Classification with Super-Prototypes and ArgumentationHamed Ayoobi, Nico Potyka, Francesca ToniAAAI 2025 · 被引用 8 次
- This Looks Like It Rather Than That: ProtoKNN For Similarity-Based ClassifiersYuki Ukai, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu FujiyoshiICLR 2023
- Interpretable Image Classification via Non-parametric Part Prototype LearningZhijie Zhu, Lei Fan, Maurice Pagnucco, Yang SongCVPR 2025
