Understanding Distributed Representations of Concepts in Deep Neural Networks without Supervision
Wonjoon Chang, Dahee Kwon, Jaesik Choi
摘要
Understanding intermediate representations of the concepts learned by deep learning classifiers is indispensable for interpreting general model behaviors. Existing approaches to reveal learned concepts often rely on human supervision, such as pre-defined concept sets or segmentation processes. In this paper, we propose a novel unsupervised method for discovering distributed representations of concepts by selecting a principal subset of neurons. Our empirical findings demonstrate that instances with similar neuron activation states tend to share coherent concepts. Based on the observations, the proposed method selects principal neurons that construct an interpretable region, namely a Relaxed Decision Region (RDR), encompassing instances with coherent concepts in the feature space. It can be utilized to identify unlabeled subclasses within data and to detect the causes of misclassifications. Furthermore, the applicability of our method across various layers discloses distinct distributed representations over the layers, which provides deeper insights into the internal mechanisms of the deep learning model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept RepresentationsDahee Kwon, Sehyun Lee, Jaesik ChoiICCV 2025
- SAEs-BrainMap: Unveiling the Emergence of Specialized Concepts in Deep Models via Brain AlignmentZiming Mao, Jia Xu, Wenxuan Pan, Mufan Xue 等ICML 2026
它引用的顶会 Paper8
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Concept Activation Regions: A Generalized Framework For Concept-Based ExplanationsJonathan Crabbé, Mihaela van der SchaarNeurIPS 2022 · 被引用 88 次
- An Efficient Explorative Sampling Considering the Generative Boundaries of Deep Generative Neural NetworksGiyoung Jeon, Haedong Jeong, Jaesik ChoiAAAI 2020 · 被引用 13 次
- ImageNet-X: Understanding Model Mistakes with Factor of Variation AnnotationsBadr Youbi Idrissi, Diane Bouchacourt, Randall Balestriero, Ivan Evtimov 等ICLR 2023 · 被引用 11 次
- Interpreting Internal Activation Patterns in Deep Temporal Neural Networks by Finding PrototypesSohee Cho, Wonjoon Chang, Ginkyeng Lee, Jaesik ChoiKDD 2021 · 被引用 10 次
相关 Paper
- Finding Representative Interpretations on Convolutional Neural NetworksPeter Cho-Ho Lam, Lingyang Chu, Maxim Torgonskiy, Jian Pei 等ICCV 2021 · 被引用 7 次
- Visual Concept Connectome (VCC): Open World Concept Discovery and Their Interlayer Connections in Deep ModelsMatthew Kowal, Richard P. Wildes, Konstantinos G. DerpanisCVPR 2024
- The Deleuzian Representation HypothesisClément Cornet, Romaric Besançon, Hervé Le BorgneICLR 2026
- ConceptExplainer: Interactive Explanation for Deep Neural Networks from a Concept PerspectiveJinbin Huang, Aditi Mishra, Bum Chul Kwon, Chris BryanIEEE VIS 2022 · 被引用 46 次
- Concept-based Explanations for Out-of-Distribution DetectorsJihye Choi, Jayaram Raghuram, Ryan Feng, Jiefeng Chen 等ICML 2023 · 被引用 18 次
