Interpretability Based Neural Network Repair
Zuohui Chen, Jun Zhou, Youcheng Sun, Jingyi Wang, Qi Xuan, Xiaoniu Yang
Abstract
Along with the prevalent use of deep neural networks (DNNs), concerns have been raised on the security threats from DNNs such as backdoors in the network. While neural network repair methods have shown to be effective for fixing the defects in DNNs, they have been also found to produce biased models, with imbalanced accuracy across different classes, or weakened adversarial robustness, allowing malicious attackers to trick the model by adding small perturbations. To address these challenges, we propose INNER, an INterpretability-based NEural Repair approach. INNER formulates the idea of neuron routing for identifying fault neurons, in which the interpretability technique model probe is used to evaluate each neuron's contribution to the undesired behaviour of the neural network. INNER then optimizes the identified neurons for repairing the neural network. We test INNER on three typical application scenarios, including backdoor attacks, adversarial attacks, and wrong predictions. Our experimental results demonstrate that INNER can effectively repair neural networks, by ensuring accuracy, fairness, and robustness. Moreover, the performance of other repair methods can be also improved by re-using the fault neurons found by INNER, justifying the generality of the proposed approach.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 35e3b24f-75e7-483b-a3b4-ce458661aa35Cited by top-tier papers3
- A Comprehensive Study of Deep Learning Model Fixing ApproachesHanmo You, Zan Wang, Zishuo Dong, Luanqi Mo et al.ICSE 2026
- Provable Repair of Deep Neural Network Defects by Preimage Synthesis and Property RefinementJianan Ma, Jingyi Wang, Qi Xuan, Zhen WangCCS 2025
- Provable Fairness Repair for Deep Neural NetworksJianan Ma, Jingyi Wang, Qi Xuan, Zhen WangASE 2025
Related papers
- Causality-Based Neural Network RepairBing Sun, Jun Sun, Long H. Pham, Tie ShiICSE 2022 · 69 citations
- Semantic-Based Neural Network RepairRichard Schumi, Jun SunISSTA 2023 · 7 citations
- AI-Lancet: Locating Error-inducing Neurons to Optimize Neural NetworksYue Zhao, Hong Zhu, Kai Chen, Shengzhi ZhangCCS 2021 · 17 citations
- Isolation-Based Debugging for Neural NetworksJialuo Chen, Jingyi Wang, Youcheng Sun, Peng Cheng et al.ISSTA 2024 · 2 citations
- Patch Synthesis for Property Repair of Deep Neural NetworksZhiming Chi, Jianan Ma, Pengfei Yang, Cheng-Chao Huang et al.ICSE 2025 · 2 citations
