Concept-Guided Interpretability via Neural Chunking
Shuchen Wu, Stephan Alaniz, Shyamgopal Karthik, Peter Dayan, Eric Schulz, Zeynep Akata
Abstract
Neural networks are often described as black boxes, reflecting the significant challenge of understanding their internal workings and interactions. We propose a different perspective that challenges the prevailing view: rather than being inscrutable, neural networks exhibit patterns in their raw population activity that mirror regularities in the training data. We refer to this as the Reflection Hypothesis and provide evidence for this phenomenon in both simple recurrent neural networks (RNNs) and complex large language models (LLMs). Building on this insight, we propose to leverage our cognitive tendency of chunking to segment high-dimensional neural population dynamics into interpretable units that reflect underlying concepts. We propose three methods to extract recurring chunks on a neural population level, complementing each other based on label availability and neural data dimensionality. Discrete sequence chunking (DSC) learns a dictionary of entities in a lower-dimensional neural space; population averaging (PA) extracts recurring entities that correspond to known labels; and unsupervised chunk discovery (UCD) can be used when labels are absent. We demonstrate the effectiveness of these methods in extracting concept-encoding entities agnostic to model architectures. These concepts can be both concrete (words), abstract (POS tags), or structural (narrative schema). Additionally, we show that extracted chunks play a causal role in network behavior, as grafting them leads to controlled and predictable changes in the model's behavior. Our work points to a new direction for interpretability, one that harnesses both cognitive principles and the structure of naturalistic data to reveal the hidden computations of complex learning systems, gradually transforming them from black boxes into systems we can begin to understand. Implementation and code are publicly available at https://github.com/swu32/Chunk-Interpretability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3fa6f30-325e-499e-bc3f-0a0b5ee0fc19Builds on16
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Revisiting Model Stitching to Compare Neural RepresentationsYamini Bansal, Preetum Nakkiran, Boaz BarakNeurIPS 2021 · 253 citations
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance EstimationThomas Fel, Victor Boutin, Louis Béthune, Rémi Cadène et al.NeurIPS 2023 · 125 citations
- On Linear Identifiability of Learned RepresentationsGeoffrey Roeder, Luke Metz, Durk KingmaICML 2021 · 107 citations
Related papers
- Decision-Guided Weighted Automata Extraction from Recurrent Neural NetworksXiyue Zhang, Xiaoning Du, Xiaofei Xie, Lei Ma et al.AAAI 2021 · 25 citations
- Towards Interpreting Recurrent Neural Networks through Probabilistic AbstractionGuoliang Dong, Jingyi Wang, Jun Sun, Yang Zhang et al.ASE 2020 · 15 citations
- LaVCa: LLM-assisted Visual Cortex CaptioningTakuya Matsuyama, Shinji Nishimoto, Yu TakagiICLR 2026 · 8 citations
- AdaAX: Explaining Recurrent Neural Networks by Learning Automata with Adaptive StatesDat Hong, Alberto Maria Segre, Tong WangKDD 2022 · 3 citations
- Learning Structure from the Ground up - Hierarchical Representation Learning by ChunkingShuchen Wu, Noémi Élteto, Ishita Dasgupta, Eric SchulzNeurIPS 2022 · 20 citations
