Understanding Distributed Representations of Concepts in Deep Neural Networks without Supervision
Wonjoon Chang, Dahee Kwon, Jaesik Choi
Abstract
Understanding intermediate representations of the concepts learned by deep learning classifiers is indispensable for interpreting general model behaviors. Existing approaches to reveal learned concepts often rely on human supervision, such as pre-defined concept sets or segmentation processes. In this paper, we propose a novel unsupervised method for discovering distributed representations of concepts by selecting a principal subset of neurons. Our empirical findings demonstrate that instances with similar neuron activation states tend to share coherent concepts. Based on the observations, the proposed method selects principal neurons that construct an interpretable region, namely a Relaxed Decision Region (RDR), encompassing instances with coherent concepts in the feature space. It can be utilized to identify unlabeled subclasses within data and to detect the causes of misclassifications. Furthermore, the applicability of our method across various layers discloses distinct distributed representations over the layers, which provides deeper insights into the internal mechanisms of the deep learning model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d311ae6d-1e88-4d36-9b6b-6e1eba6b17f7Cited by top-tier papers2
- Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept RepresentationsDahee Kwon, Sehyun Lee, Jaesik ChoiICCV 2025
- SAEs-BrainMap: Unveiling the Emergence of Specialized Concepts in Deep Models via Brain AlignmentZiming Mao, Jia Xu, Wenxuan Pan, Mufan Xue et al.ICML 2026
Builds on8
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Concept Activation Regions: A Generalized Framework For Concept-Based ExplanationsJonathan Crabbé, Mihaela van der SchaarNeurIPS 2022 · 88 citations
- An Efficient Explorative Sampling Considering the Generative Boundaries of Deep Generative Neural NetworksGiyoung Jeon, Haedong Jeong, Jaesik ChoiAAAI 2020 · 13 citations
- ImageNet-X: Understanding Model Mistakes with Factor of Variation AnnotationsBadr Youbi Idrissi, Diane Bouchacourt, Randall Balestriero, Ivan Evtimov et al.ICLR 2023 · 11 citations
- Interpreting Internal Activation Patterns in Deep Temporal Neural Networks by Finding PrototypesSohee Cho, Wonjoon Chang, Ginkyeng Lee, Jaesik ChoiKDD 2021 · 10 citations
Related papers
- Finding Representative Interpretations on Convolutional Neural NetworksPeter Cho-Ho Lam, Lingyang Chu, Maxim Torgonskiy, Jian Pei et al.ICCV 2021 · 7 citations
- Visual Concept Connectome (VCC): Open World Concept Discovery and Their Interlayer Connections in Deep ModelsMatthew Kowal, Richard P. Wildes, Konstantinos G. DerpanisCVPR 2024
- The Deleuzian Representation HypothesisClément Cornet, Romaric Besançon, Hervé Le BorgneICLR 2026
- ConceptExplainer: Interactive Explanation for Deep Neural Networks from a Concept PerspectiveJinbin Huang, Aditi Mishra, Bum Chul Kwon, Chris BryanIEEE VIS 2022 · 46 citations
- Concept-based Explanations for Out-of-Distribution DetectorsJihye Choi, Jayaram Raghuram, Ryan Feng, Jiefeng Chen et al.ICML 2023 · 18 citations
