Interpretability Based Neural Network Repair
Zuohui Chen, Jun Zhou, Youcheng Sun, Jingyi Wang, Qi Xuan, Xiaoniu Yang
摘要
Along with the prevalent use of deep neural networks (DNNs), concerns have been raised on the security threats from DNNs such as backdoors in the network. While neural network repair methods have shown to be effective for fixing the defects in DNNs, they have been also found to produce biased models, with imbalanced accuracy across different classes, or weakened adversarial robustness, allowing malicious attackers to trick the model by adding small perturbations. To address these challenges, we propose INNER, an INterpretability-based NEural Repair approach. INNER formulates the idea of neuron routing for identifying fault neurons, in which the interpretability technique model probe is used to evaluate each neuron's contribution to the undesired behaviour of the neural network. INNER then optimizes the identified neurons for repairing the neural network. We test INNER on three typical application scenarios, including backdoor attacks, adversarial attacks, and wrong predictions. Our experimental results demonstrate that INNER can effectively repair neural networks, by ensuring accuracy, fairness, and robustness. Moreover, the performance of other repair methods can be also improved by re-using the fault neurons found by INNER, justifying the generality of the proposed approach.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- A Comprehensive Study of Deep Learning Model Fixing ApproachesHanmo You, Zan Wang, Zishuo Dong, Luanqi Mo 等ICSE 2026
- Provable Repair of Deep Neural Network Defects by Preimage Synthesis and Property RefinementJianan Ma, Jingyi Wang, Qi Xuan, Zhen WangCCS 2025
- Provable Fairness Repair for Deep Neural NetworksJianan Ma, Jingyi Wang, Qi Xuan, Zhen WangASE 2025
相关 Paper
- Causality-Based Neural Network RepairBing Sun, Jun Sun, Long H. Pham, Tie ShiICSE 2022 · 被引用 69 次
- Semantic-Based Neural Network RepairRichard Schumi, Jun SunISSTA 2023 · 被引用 7 次
- AI-Lancet: Locating Error-inducing Neurons to Optimize Neural NetworksYue Zhao, Hong Zhu, Kai Chen, Shengzhi ZhangCCS 2021 · 被引用 17 次
- Isolation-Based Debugging for Neural NetworksJialuo Chen, Jingyi Wang, Youcheng Sun, Peng Cheng 等ISSTA 2024 · 被引用 2 次
- Patch Synthesis for Property Repair of Deep Neural NetworksZhiming Chi, Jianan Ma, Pengfei Yang, Cheng-Chao Huang 等ICSE 2025 · 被引用 2 次
