Model Orthogonalization: Class Distance Hardening in Neural Networks for Better Security
Guanhong Tao, Yingqi Liu, Guangyu Shen, Qiuling Xu, Shengwei An, Zhuo Zhang, Xiangyu Zhang
Abstract
The distance between two classes for a deep learning classifier can be measured by the level of difficulty in flipping all (or majority of) samples in a class to the other. The class distances of many pre-trained models in the wild are very small and do not align well with humans’ intuition (e.g., classes turtle and bird have smaller distance than classes cat and dog), making the models vulnerable to backdoor attacks, which aim to cause misclassification by stamping a specific pattern to inputs. We propose a novel model hardening technique called model orthogonalization which is an add-on training step to pretrained models, including clean models, poisoned models, and adversarially trained models. It can substantially enlarge class distances with reasonable training cost and without much accuracy degradation. Our evaluation on 5 datasets with 22 model structures show that our technique can enlarge class distances by 177.63% on average with less than 1% accuracy loss, outperforming existing hardening techniques such as adversarial training, universal adversarial perturbation, and directly using generated backdoors. It reduces 80% false positives for a state-of-the-art backdoor scanner as the enlarged class distances allow the scanner to easily distinguish clean and poisoned models, and substantially outperforms three existing techniques in removing injected backdoors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cdd345af-e88e-47af-9103-2d9cecd3ff49Cited by top-tier papers28
- BppAttack: Stealthy and Efficient Trojan Attacks against Deep Neural Networks via Image Quantization and Contrastive Adversarial LearningZhenting Wang, Juan Zhai, Shiqing MaCVPR 2022 · 92 citations
- Black-box Backdoor Defense via Zero-shot Image PurificationYucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan et al.NeurIPS 2023 · 66 citations
- Constrained Optimization with Dynamic Bound-scaling for Effective NLP Backdoor DefenseGuangyu Shen, Yingqi Liu, Guanhong Tao, Qiuling Xu et al.ICML 2022 · 58 citations
- Robust Backdoor Detection for Deep Learning via Topological Evolution DynamicsXiaoxing Mo, Yechao Zhang, Leo Yu Zhang, Wei Luo et al.S&P 2024 · 39 citations
- Distribution Preserving Backdoor Attack in Self-supervised LearningGuanhong Tao, Zhenting Wang, Shiwei Feng, Guangyu Shen et al.S&P 2024 · 32 citations
Builds on29
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- DBA: Distributed Backdoor Attacks against Federated LearningChulin Xie, Keli Huang, Pin-Yu Chen, Bo LiICLR 2020 · 901 citations
- Attack of the Tails: Yes, You Really Can Backdoor Federated LearningHongyi Wang, Kartik Sreenivasan, Shashank Rajput, Harit Vishwakarma et al.NeurIPS 2020 · 862 citations
Related papers
- Improving the Sensitivity of Backdoor Detectors via Class Subspace OrthogonalizationGuangmingmei Yang, David Miller, George KesidisICML 2026 · 3 citations
- Universal Backdoor AttacksBenjamin Schneider, Nils Lukas, Florian KerschbaumICLR 2024 · 17 citations
- BEAGLE: Forensics of Deep Learning Backdoor Attack for Better DefenseSiyuan Cheng, Guanhong Tao, Yingqi Liu, Shengwei An et al.NDSS 2023
- Backdoor Mitigation by Distance-Driven DetoxificationShaokui Wei, Jiayin Liu, Hongyuan ZhaICCV 2025 · 1 citation
- MEDIC: Remove Model Backdoors via Importance Driven CloningQiuling Xu, Guanhong Tao, Jean Honorio, Yingqi Liu et al.CVPR 2023
