Adversarial Robustness through Random Weight Sampling
Yanxiang Ma, Minjing Dong, Chang Xu
Abstract
Deep neural networks have been found to be vulnerable in a variety of tasks. Ad-versarial attacks can manipulate network outputs, resulting in incorrect predictions. Adversarial defense methods aim to improve the adversarial robustness of networks by countering potential attacks. In addition to traditional defense approaches, randomized defense mechanisms have recently received increasing attention from researchers. These methods introduce different types of perturbations during the inference phase to destabilize adversarial attacks. Although promising empirical re-sults have been demonstrated by these approaches, the defense performance is quite sensitive to the randomness parameters, which are always manually tuned without further analysis. On the contrary, we propose incorporating random weights into the optimization to exploit the potential of randomized defense fully. To perform better optimization of randomness parameters, we conduct a theoretical analysis of the connections between randomness parameters and gradient similarity as well as natural performance. From these two aspects, we suggest imposing theoretically-guided constraints on random weights during optimizations, as these weights play a critical role in balancing natural performance and adversarial robustness. We derive both the upper and lower bounds of random weight parameters by considering prediction bias and gradient similarity. In this study, we introduce the Constrained Trainable Random Weight (CTRW), which adds random weight parameters to the optimization and includes a constraint guided by the upper and lower bounds to achieve better trade-offs between natural and robust accuracy. We evaluate the effectiveness of CTRW on several datasets and benchmark convolutional neural networks. Our results indicate that our model achieves a robust accuracy approximately 16% to 17% higher than the baseline model under PGD-20 and 22% to 25% higher on Auto Attack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b87e5c27-b8d4-4c46-852a-149f1d59da60Cited by top-tier papers10
- Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language ModelsLu Yu, Haiyang Zhang, Changsheng XuNeurIPS 2024 · 29 citations
- A Hard-Label Black-Box Evasion Attack against ML-based Malicious Traffic Detection SystemsZixuan Liu, Yi Zhao, Zhuotao Liu, Qi Li et al.NDSS 2026 · 3 citations
- Effective and Robust Multimodal Medical Image AnalysisJoy Dhar, Nayyar Zaidi, Maryam HaghighatKDD 2026 · 1 citation
- First Line of Defense: A Robust First Layer Mitigates Adversarial AttacksJanani Suresh, Nancy Nayak, Sheetal KalyaniAAAI 2025 · 1 citation
- Adversarial Robustness via Deformable Convolution with StochasticityYanxiang Ma, Zixuan Huang, Minjing Dong, Shan You et al.ICML 2025
Builds on13
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Prediction Poisoning: Towards Defenses Against DNN Model Stealing AttacksTribhuvanesh Orekondy, Bernt Schiele, Mario FritzICLR 2020 · 194 citations
- Dataset Distillation via FactorizationSonghua Liu, Kai Wang, Xingyi Yang, Jingwen Ye et al.NeurIPS 2022 · 190 citations
Related papers
- Randomized Adversarial Training via Taylor ExpansionGaojie Jin, Xinping Yi, Dengyu Wu, Ronghui Mu et al.CVPR 2023
- Adversarial Robustness via Random Projection FiltersMinjing Dong, Chang XuCVPR 2023
- Constrained Gradient Descent: A Powerful and Principled Evasion Attack Against Neural NetworksWeiran Lin, Keane Lucas, Lujo Bauer, Michael K. Reiter et al.ICML 2022 · 5 citations
- Random Entangled Tokens for Adversarially Robust Vision TransformerHuihui Gong, Minjing Dong, Siqi Ma, Seyit Camtepe et al.CVPR 2024
- Enhancing Adversarial Training with Second-Order Statistics of WeightsGaojie Jin, Xinping Yi, Wei Huang, Sven Schewe et al.CVPR 2022 · 46 citations
