Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis
Arpita Chowdhury, Dipanjyoti Paul, Zheda Mai, Jianyang Gu, Ziheng Zhang, Kazi Sajeed Mehrab, Elizabeth G. Campolongo, Daniel I. Rubenstein, Charles V. Stewart, Anuj Karpatne, Tanya Y. Berger-Wolf, Yu Su, Wei-Lun Chao
Abstract
We present a simple approach to make pre-trained Vision Transformers (ViTs) interpretable for fine-grained analysis, aiming to identify and localize the traits that distinguish visually similar categories, such as bird species. Pre-trained ViTs, such as DINO, have demonstrated remarkable capabilities in extracting localized, discriminative features. However, saliency maps like Grad-CAM often fail to identify these traits, producing blurred, coarse heatmaps that highlight entire objects instead. We propose a novel approach, Prompt Class Attention Map (Prompt-CAM), to address this limitation. Prompt-CAM learns class-specific prompts for a pre-trained ViT and uses the corresponding outputs for classification. To correctly classify an image, the true-class prompt must attend to unique image patches not present in other classes' images (i.e., traits). As a result, the true class's multi-head attention maps reveal traits and their locations. Implementation-wise, Prompt-CAM is almost a "free lunch," requiring only a modification to the prediction head of Visual Prompt Tuning (VPT). This makes Prompt-CAM easy to train and apply, in stark contrast to other interpretable methods that require designing specific models and training processes. Extensive empirical studies on a dozen datasets from various domains (e.g., birds, fishes, insects, fungi, flowers, food, and cars) validate the superior interpretation capability of Prompt-CAM. The source code and demo are available at https://github.com/Imageomics/Prompt_CAM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 682462d2-d3b8-4294-a4ef-720ad364eb91Cited by top-tier papers3
- Visual Instance-aware Prompt TuningXi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang et al.ACM MM 2025 · 12 citations
- AVION: Aerial Vision–Language Instruction from Offline Teacher to Prompt-Tuned NetworkYu Hu, Jianyang Gu, Hao Liu, Yue Cao et al.CVPR 2026
- Towards Robust Vision Transformers: Path Dependency Analysis and a Simple Two-Stage Adversarial TrainingSeongmin Kim, Byung Cheol SongCVPR 2026
Builds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
Related papers
- Token Coordinated Prompt Attention is Needed for Visual PromptingZichen Liu, Xu Zou, Gang Hua, Jiahuan ZhouICML 2025
- Fair-VPT: Fair Visual Prompt Tuning for Image ClassificationSungho Park, Hyeran ByunCVPR 2024 · 12 citations
- Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-AttentionSaebom Leem, Hyunseok SeoAAAI 2024 · 40 citations
- Less Attention is More: Prompt Transformer for Generalized Category DiscoveryWei Zhang, Baopeng Zhang, Zhu Teng, Wenxin Luo et al.CVPR 2025
- Improving Visual Prompt Tuning for Self-supervised Vision TransformersSeungryong Yoo, Eunji Kim, Dahuin Jung, Jungbeom Lee et al.ICML 2023 · 74 citations
