NeurFlow: Interpreting Neural Networks through Neuron Groups and Functional Interactions
Tue Minh Cao, Nhat Hoang-Xuan, Hieu H. Pham, Phi Le Nguyen, My T. Thai
Abstract
Understanding the inner workings of neural networks is essential for enhancing model performance and interpretability. Current research predominantly focuses on examining the connection between individual neurons and the model's final predictions, which suffers from challenges in interpreting the internal workings of the model, particularly when neurons encode multiple unrelated features. In this paper, we propose a novel framework that transitions the focus from analyzing individual neurons to investigating groups of neurons, shifting the emphasis from neuron-output relationships to the functional interactions between neurons. Our automated framework, NeurFlow, first identifies core neurons and clusters them into groups based on shared functional relationships, enabling a more coherent and interpretable view of the network's internal processes. This approach facilitates the construction of a hierarchical circuit representing neuron interactions across layers, thus improving interpretability while reducing computational costs. Our extensive empirical studies validate the fidelity of our proposed NeurFlow. Additionally, we showcase its utility in practical applications such as image debugging and automatic concept labeling, thereby highlighting its potential to advance the field of neural network explainability. 4
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Constructing Interpretable Features from Compositional Neuron GroupsOr David Shafran, Atticus Geiger, Mor GevaACL 2026 · 4 citations
- Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPsTanya Chowdhury, Atharva Nijasure, Yair Zick, James AllanICLR 2026 · 1 citation
- CL-Guard: Defending DNNs Against Backdoors via Fine-Grained Neuron Analysis and Collaborative Dual-Network LearningJie Xiao, Yuhao Huang, Yanjiao Gao, Aizhu Liu et al.AAAI 2026
Builds on17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- Neuron Shapley: Discovering the Responsible NeuronsAmirata Ghorbani, James Y. ZouNeurIPS 2020 · 160 citations
- Invertible Concept-based Explanations for CNN Models with Non-negative Concept Activation VectorsRuihan Zhang, Prashan Madumal, Tim Miller, Krista A. Ehinger et al.AAAI 2021 · 140 citations
Related papers
- Knowledge-Aware Neuron Interpretation for Scene ClassificationYong Guan, Freddy Lécué, Jiaoyan Chen, Ru Li et al.AAAI 2024 · 3 citations
- NeuroCartography: Scalable Automatic Visual Summarization of Concepts in Deep Neural NetworksHaekyu Park, Nilaksh Das, Rahul Duggal, Austin P. Wright et al.IEEE VIS 2021 · 28 citations
- Inner Information Analysis Algorithm for Deep Neural Network based on CommunityGuipeng Lan, Shuai Xiao, Meng Xi, Jiabao Wen et al.ICLR 2025
- Select, Hypothesize and Verify: Towards Verified Neuron Concept InterpretationZeBin Ji, Yang Hu, Xiuli Bi, Bo Liu et al.CVPR 2026
- Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept RepresentationsDahee Kwon, Sehyun Lee, Jaesik ChoiICCV 2025
