A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis
Dipanjyoti Paul, Arpita Chowdhury, Xinqi Xiong, Feng-Ju Chang, David Edward Carlyn, Samuel Stevens, Kaiya Provost, Anuj Karpatne, Bryan Carstens, Daniel I. Rubenstein, Charles V. Stewart, Tanya Y. Berger-Wolf
摘要
We present a novel usage of Transformers to make image classification interpretable. Unlike mainstream classifiers that wait until the last fully connected layer to incorporate class information to make predictions, we investigate a proactive approach, asking each class to search for itself in an image. We realize this idea via a Transformer encoder-decoder inspired by DEtection TRansformer (DETR). We learn"class-specific"queries (one for each class) as input to the decoder, enabling each class to localize its patterns in an image via cross-attention. We name our approach INterpretable TRansformer (INTR), which is fairly easy to implement and exhibits several compelling properties. We show that INTR intrinsically encourages each class to attend distinctively; the cross-attention weights thus provide a faithful interpretation of the prediction. Interestingly, via"multi-head"cross-attention, INTR could identify different"attributes"of a class, making it particularly suitable for fine-grained classification and analysis, which we demonstrate on eight datasets. Our code and pre-trained models are publicly accessible at the Imageomics Institute GitHub site: https://github.com/Imageomics/INTR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Interpretable Image Classification with Adaptive Prototype-based Vision TransformersChiyu Ma, Jon Donnelly, Wenjun Liu, Soroush Vosoughi 等NeurIPS 2024 · 被引用 48 次
- AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation ModelsZheda Mai, Arpita Chowdhury, Zihe Wang, Sooyoung Jeon 等CVPR 2026 · 被引用 7 次
- Large Connectome Model: An fMRI Foundation Model of Brain Connectomes Empowered by Brain-Environment Interaction in Multitask Learning LandscapeZiquan Wei, Tingting Dan, Guorong WuAAAI 2026 · 被引用 2 次
- Finer-CAM: Spotting the Difference Reveals Finer Details for Visual ExplanationZiheng Zhang, Jianyang Gu, Arpita Chowdhury, Zheda Mai 等CVPR 2025
- What Do You See in Common? Learning Hierarchical Prototypes over Tree-of-Life to Discover Evolutionary TraitsHarish Babu Manogaran, M. Maruf, Arka Daw, Kazi Sajeed Mehrab 等ICLR 2025
它引用的顶会 Paper13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
相关 Paper
- Analyzing Vision Transformers for Image Classification in Class Embedding SpaceMartina G. Vilas, Timothy Schaumlöffel, Gemma RoigNeurIPS 2023 · 被引用 43 次
- DECO: Unleashing the Potential of ConvNets for Query-based Detection and SegmentationXinghao Chen, Siwei Li, Yijing Yang, Yunhe WangICLR 2025
- DESTR: Object Detection with Split TransformerLiqiang He, Sinisa TodorovicCVPR 2022 · 被引用 63 次
- ClassFormer: Exploring Class-Aware Dependency with Transformer for Medical Image SegmentationHuimin Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong 等AAAI 2023 · 被引用 5 次
- PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object DetectionZhengjian Kang, Jun Zhuang, Kangtong Mo, Qi Chen 等CVPR 2026 · 被引用 6 次
