Explaining Deep Convolutional Neural Networks via Latent Visual-Semantic Filter Attention
Yu Yang, Seungbae Kim, Jungseock Joo
摘要
Interpretability is an important property for visual mod-els as it helps researchers and users understand the in-ternal mechanism of a complex model. However, gener-ating semantic explanations about the learned representation is challenging without direct supervision to produce such explanations. We propose a general framework, La-tent Visual Semantic Explainer (LaViSE), to teach any ex-isting convolutional neural network to generate text de-scriptions about its own latent representations at the filter level. Our method constructs a mapping between the vi-sual and semantic spaces using generic image datasets, using images and category names. It then transfers the map-ping to the target domain which does not have semantic la-bels. The proposedframework employs a modular structure and enables to analyze any trained network whether or not its original training data is available. We show that our method can generate novel descriptions for learned filters beyond the set of categories defined in the training dataset and perform an extensive evaluation on multiple datasets. We also demonstrate a novel application of our method for unsupervised dataset bias analysis which allows us to auto-matically discover hidden biases in datasets or compare dif-ferent subsets without using additional labels. The dataset and code are made public to facilitate further research. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> https://github.com/YuYang0901/LaViSE
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Mitigating Spurious Correlations in Multi-modal Models during Fine-tuningYu Yang, Besmira Nushi, Hamid Palangi, Baharan MirzasoleimanICML 2023 · 被引用 65 次
- PICNN: A Pathway towards Interpretable Convolutional Neural NetworksWengang Guo, Jiayi Yang, Huilin Yin, Qijun Chen 等AAAI 2024 · 被引用 6 次
- ConceptScope: Characterizing Dataset Bias via Disentangled Visual ConceptsJinho Choi, Hyesu Lim, Steffen Schneider, Jaegul ChooNeurIPS 2025 · 被引用 5 次
- Game on Tree: Visual Hallucination Mitigation via Coarse-to-Fine View Tree and Game TheoryXianwei Zhuang, Zhihong Zhu, Zhanpeng Chen, Yuxin Xie 等EMNLP 2024 · 被引用 1 次
- Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation LearningPeng Jin, Jinfa Huang, Pengfei Xiong, Shangxuan Tian 等CVPR 2023
它引用的顶会 Paper11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- nocaps: novel object captioning at scaleHarsh Agrawal, Peter Anderson, Karan Desai, Yufei Wang 等ICCV 2019 · 被引用 631 次
- Understanding and Mitigating Annotation Bias in Facial Expression RecognitionYunliang Chen, Jungseock JooICCV 2021 · 被引用 108 次
- Neural Prototype Trees for Interpretable Fine-Grained Image RecognitionMeike Nauta, Ron van Bree, Christin SeifertCVPR 2021
- GANmut: Learning Interpretable Conditional Space for Gamut of EmotionsStefano d'Apolito, Danda Pani Paudel, Zhiwu Huang, Andrés Romero 等CVPR 2021
相关 Paper
- Interpretable Generative Adversarial NetworksChao Li, Kelu Yao, Jin Wang, Boyu Diao 等AAAI 2022 · 被引用 19 次
- A Disentangling Invertible Interpretation Network for Explaining Latent RepresentationsPatrick Esser, Robin Rombach, Björn OmmerCVPR 2020
- ConceptExplainer: Interactive Explanation for Deep Neural Networks from a Concept PerspectiveJinbin Huang, Aditi Mishra, Bum Chul Kwon, Chris BryanIEEE VIS 2022 · 被引用 46 次
- CoSy: Evaluating Textual Explanations of NeuronsLaura Kopf, Philine Lou Bommer, Anna Hedström, Sebastian Lapuschkin 等NeurIPS 2024 · 被引用 23 次
- CNN Explainer: Learning Convolutional Neural Networks with Interactive VisualizationZijie J. Wang, Robert Turko, Omar Shaikh, Haekyu Park 等IEEE VIS 2020 · 被引用 341 次
