NeurFlow: Interpreting Neural Networks through Neuron Groups and Functional Interactions
Tue Minh Cao, Nhat Hoang-Xuan, Hieu H. Pham, Phi Le Nguyen, My T. Thai
摘要
Understanding the inner workings of neural networks is essential for enhancing model performance and interpretability. Current research predominantly focuses on examining the connection between individual neurons and the model's final predictions, which suffers from challenges in interpreting the internal workings of the model, particularly when neurons encode multiple unrelated features. In this paper, we propose a novel framework that transitions the focus from analyzing individual neurons to investigating groups of neurons, shifting the emphasis from neuron-output relationships to the functional interactions between neurons. Our automated framework, NeurFlow, first identifies core neurons and clusters them into groups based on shared functional relationships, enabling a more coherent and interpretable view of the network's internal processes. This approach facilitates the construction of a hierarchical circuit representing neuron interactions across layers, thus improving interpretability while reducing computational costs. Our extensive empirical studies validate the fidelity of our proposed NeurFlow. Additionally, we showcase its utility in practical applications such as image debugging and automatic concept labeling, thereby highlighting its potential to advance the field of neural network explainability. 4
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Constructing Interpretable Features from Compositional Neuron GroupsOr David Shafran, Atticus Geiger, Mor GevaACL 2026 · 被引用 4 次
- Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPsTanya Chowdhury, Atharva Nijasure, Yair Zick, James AllanICLR 2026 · 被引用 1 次
- CL-Guard: Defending DNNs Against Backdoors via Fine-Grained Neuron Analysis and Collaborative Dual-Network LearningJie Xiao, Yuhao Huang, Yanjiao Gao, Aizhu Liu 等AAAI 2026
它引用的顶会 Paper17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 被引用 229 次
- Neuron Shapley: Discovering the Responsible NeuronsAmirata Ghorbani, James Y. ZouNeurIPS 2020 · 被引用 160 次
- Invertible Concept-based Explanations for CNN Models with Non-negative Concept Activation VectorsRuihan Zhang, Prashan Madumal, Tim Miller, Krista A. Ehinger 等AAAI 2021 · 被引用 140 次
相关 Paper
- Knowledge-Aware Neuron Interpretation for Scene ClassificationYong Guan, Freddy Lécué, Jiaoyan Chen, Ru Li 等AAAI 2024 · 被引用 3 次
- NeuroCartography: Scalable Automatic Visual Summarization of Concepts in Deep Neural NetworksHaekyu Park, Nilaksh Das, Rahul Duggal, Austin P. Wright 等IEEE VIS 2021 · 被引用 28 次
- Inner Information Analysis Algorithm for Deep Neural Network based on CommunityGuipeng Lan, Shuai Xiao, Meng Xi, Jiabao Wen 等ICLR 2025
- Select, Hypothesize and Verify: Towards Verified Neuron Concept InterpretationZeBin Ji, Yang Hu, Xiuli Bi, Bo Liu 等CVPR 2026
- Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept RepresentationsDahee Kwon, Sehyun Lee, Jaesik ChoiICCV 2025
