Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
Ziheng Zhang, Jianyang Gu, Arpita Chowdhury, Zheda Mai, David Carlyn, Tanya Y. Berger-Wolf, Yu Su, Wei-Lun Chao
Abstract
Class activation map (CAM) has been widely used to highlight image regions that contribute to class predictions. Despite its simplicity and computational efficiency, CAM often struggles to identify discriminative regions that distinguish visually similar fine-grained classes. Prior efforts address this limitation by introducing more sophisticated explanation processes, but at the cost of extra complexity. In this paper, we propose Finer-CAM, a method that retains CAM's efficiency while achieving precise localization of discriminative regions. Our key insight is that the deficiency of CAM lies not in "how" it explains, but in "what" it explains. Specifically, previous methods attempt to identify all cues contributing to the target class's logit value, which inadvertently also activates regions predictive of visually similar classes. By explicitly comparing the target class with similar classes and spotting their differences, Finer-CAM suppresses features shared with other classes and empha-
sizes the unique, discriminative details of the target class. Finer-CAM is easy to implement, compatible with various CAM methods, and can be extended to multi-modal models for accurate localization of specific concepts. Additionally, Finer-CAM allows adjustable comparison strength, enabling users to selectively highlight coarse object contours or fine discriminative details. Quantitatively, we show that masking out the top 5% of activated pixels by Finer-CAM results in a larger relative confidence drop compared to baselines. The source code and demo are available at https://github.com/Imageomics/Finer-CAM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41398b75-a7cb-4b91-9368-28f69bbb7802Cited by top-tier papers6
- AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation ModelsZheda Mai, Arpita Chowdhury, Zihe Wang, Sooyoung Jeon et al.CVPR 2026 · 7 citations
- BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation ModelsZiheng Zhang, Xinyue Ma, Arpita Chowdhury, Elizabeth G Campolongo et al.ICLR 2026 · 3 citations
- MaskDiME: Adaptive Masked Diffusion for Precise and Efficient Visual Counterfactual ExplanationsChanglu Guo, Anders Nymark Christensen, Anders Bjorholm Dahl, Morten Rieger HannemoseCVPR 2026 · 2 citations
- Diffusion-CAM: Faithful Visual Explanations for dMLLMsHaomin Zuo, Yidi Li, Luoxiao Yang, Xiaofeng ZhangACL 2026
- MedLIME: A Distribution-Aligned and Evidence-Supported Framework for Medical Saliency ExplanationsRaghav Magazine, Xingjian Li, Min XuCVPR 2026
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu et al.ICML 2020 · 148 citations
- A Simple Interpretable Transformer for Fine-Grained Image Classification and AnalysisDipanjyoti Paul, Arpita Chowdhury, Xinqi Xiong, Feng-Ju Chang et al.ICLR 2024 · 27 citations
Related papers
- Empowering CAM-Based Methods with Capability to Generate Fine-Grained and High-Faithfulness ExplanationsChangqing Qiu, Fusheng Jin, Yining ZhangAAAI 2024 · 11 citations
- Keep CALM and Improve Visual Feature AttributionJae-Myung Kim, Junsuk Choe, Zeynep Akata, Seong Joon OhICCV 2021 · 22 citations
- ODAM: Gradient-based Instance-Specific Visual Explanations for Object DetectionChenyang Zhao, Antoni B. ChanICLR 2023 · 4 citations
- Bridging the Gap between Classification and Localization for Weakly Supervised Object LocalizationEunji Kim, Siwon Kim, Jungbeom Lee, Hyunwoo Kim et al.CVPR 2022 · 44 citations
- Partial Class Activation Attention for Semantic SegmentationSun'ao Liu, Hongtao Xie, Hai Xu, Yongdong Zhang et al.CVPR 2022 · 47 citations
