Explaining Deep Convolutional Neural Networks via Latent Visual-Semantic Filter Attention
Yu Yang, Seungbae Kim, Jungseock Joo
Abstract
Interpretability is an important property for visual mod-els as it helps researchers and users understand the in-ternal mechanism of a complex model. However, gener-ating semantic explanations about the learned representation is challenging without direct supervision to produce such explanations. We propose a general framework, La-tent Visual Semantic Explainer (LaViSE), to teach any ex-isting convolutional neural network to generate text de-scriptions about its own latent representations at the filter level. Our method constructs a mapping between the vi-sual and semantic spaces using generic image datasets, using images and category names. It then transfers the map-ping to the target domain which does not have semantic la-bels. The proposedframework employs a modular structure and enables to analyze any trained network whether or not its original training data is available. We show that our method can generate novel descriptions for learned filters beyond the set of categories defined in the training dataset and perform an extensive evaluation on multiple datasets. We also demonstrate a novel application of our method for unsupervised dataset bias analysis which allows us to auto-matically discover hidden biases in datasets or compare dif-ferent subsets without using additional labels. The dataset and code are made public to facilitate further research. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> https://github.com/YuYang0901/LaViSE
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 63fdd5fe-ac71-4ef2-8727-90c2656ef256Cited by top-tier papers5
- Mitigating Spurious Correlations in Multi-modal Models during Fine-tuningYu Yang, Besmira Nushi, Hamid Palangi, Baharan MirzasoleimanICML 2023 · 65 citations
- PICNN: A Pathway towards Interpretable Convolutional Neural NetworksWengang Guo, Jiayi Yang, Huilin Yin, Qijun Chen et al.AAAI 2024 · 6 citations
- ConceptScope: Characterizing Dataset Bias via Disentangled Visual ConceptsJinho Choi, Hyesu Lim, Steffen Schneider, Jaegul ChooNeurIPS 2025 · 5 citations
- Game on Tree: Visual Hallucination Mitigation via Coarse-to-Fine View Tree and Game TheoryXianwei Zhuang, Zhihong Zhu, Zhanpeng Chen, Yuxin Xie et al.EMNLP 2024 · 1 citation
- Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation LearningPeng Jin, Jinfa Huang, Pengfei Xiong, Shangxuan Tian et al.CVPR 2023
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- nocaps: novel object captioning at scaleHarsh Agrawal, Peter Anderson, Karan Desai, Yufei Wang et al.ICCV 2019 · 631 citations
- Understanding and Mitigating Annotation Bias in Facial Expression RecognitionYunliang Chen, Jungseock JooICCV 2021 · 108 citations
- Neural Prototype Trees for Interpretable Fine-Grained Image RecognitionMeike Nauta, Ron van Bree, Christin SeifertCVPR 2021
- GANmut: Learning Interpretable Conditional Space for Gamut of EmotionsStefano d'Apolito, Danda Pani Paudel, Zhiwu Huang, Andrés Romero et al.CVPR 2021
Related papers
- Interpretable Generative Adversarial NetworksChao Li, Kelu Yao, Jin Wang, Boyu Diao et al.AAAI 2022 · 19 citations
- A Disentangling Invertible Interpretation Network for Explaining Latent RepresentationsPatrick Esser, Robin Rombach, Björn OmmerCVPR 2020
- ConceptExplainer: Interactive Explanation for Deep Neural Networks from a Concept PerspectiveJinbin Huang, Aditi Mishra, Bum Chul Kwon, Chris BryanIEEE VIS 2022 · 46 citations
- CoSy: Evaluating Textual Explanations of NeuronsLaura Kopf, Philine Lou Bommer, Anna Hedström, Sebastian Lapuschkin et al.NeurIPS 2024 · 23 citations
- CNN Explainer: Learning Convolutional Neural Networks with Interactive VisualizationZijie J. Wang, Robert Turko, Omar Shaikh, Haekyu Park et al.IEEE VIS 2020 · 341 citations
