Adversarial Robustness through Random Weight Sampling
Yanxiang Ma, Minjing Dong, Chang Xu
摘要
Deep neural networks have been found to be vulnerable in a variety of tasks. Ad-versarial attacks can manipulate network outputs, resulting in incorrect predictions. Adversarial defense methods aim to improve the adversarial robustness of networks by countering potential attacks. In addition to traditional defense approaches, randomized defense mechanisms have recently received increasing attention from researchers. These methods introduce different types of perturbations during the inference phase to destabilize adversarial attacks. Although promising empirical re-sults have been demonstrated by these approaches, the defense performance is quite sensitive to the randomness parameters, which are always manually tuned without further analysis. On the contrary, we propose incorporating random weights into the optimization to exploit the potential of randomized defense fully. To perform better optimization of randomness parameters, we conduct a theoretical analysis of the connections between randomness parameters and gradient similarity as well as natural performance. From these two aspects, we suggest imposing theoretically-guided constraints on random weights during optimizations, as these weights play a critical role in balancing natural performance and adversarial robustness. We derive both the upper and lower bounds of random weight parameters by considering prediction bias and gradient similarity. In this study, we introduce the Constrained Trainable Random Weight (CTRW), which adds random weight parameters to the optimization and includes a constraint guided by the upper and lower bounds to achieve better trade-offs between natural and robust accuracy. We evaluate the effectiveness of CTRW on several datasets and benchmark convolutional neural networks. Our results indicate that our model achieves a robust accuracy approximately 16% to 17% higher than the baseline model under PGD-20 and 22% to 25% higher on Auto Attack.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language ModelsLu Yu, Haiyang Zhang, Changsheng XuNeurIPS 2024 · 被引用 29 次
- A Hard-Label Black-Box Evasion Attack against ML-based Malicious Traffic Detection SystemsZixuan Liu, Yi Zhao, Zhuotao Liu, Qi Li 等NDSS 2026 · 被引用 3 次
- Effective and Robust Multimodal Medical Image AnalysisJoy Dhar, Nayyar Zaidi, Maryam HaghighatKDD 2026 · 被引用 1 次
- First Line of Defense: A Robust First Layer Mitigates Adversarial AttacksJanani Suresh, Nancy Nayak, Sheetal KalyaniAAAI 2025 · 被引用 1 次
- Adversarial Robustness via Deformable Convolution with StochasticityYanxiang Ma, Zixuan Huang, Minjing Dong, Shan You 等ICML 2025
它引用的顶会 Paper13
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
- Prediction Poisoning: Towards Defenses Against DNN Model Stealing AttacksTribhuvanesh Orekondy, Bernt Schiele, Mario FritzICLR 2020 · 被引用 194 次
- Dataset Distillation via FactorizationSonghua Liu, Kai Wang, Xingyi Yang, Jingwen Ye 等NeurIPS 2022 · 被引用 190 次
相关 Paper
- Randomized Adversarial Training via Taylor ExpansionGaojie Jin, Xinping Yi, Dengyu Wu, Ronghui Mu 等CVPR 2023
- Adversarial Robustness via Random Projection FiltersMinjing Dong, Chang XuCVPR 2023
- Constrained Gradient Descent: A Powerful and Principled Evasion Attack Against Neural NetworksWeiran Lin, Keane Lucas, Lujo Bauer, Michael K. Reiter 等ICML 2022 · 被引用 5 次
- Random Entangled Tokens for Adversarially Robust Vision TransformerHuihui Gong, Minjing Dong, Siqi Ma, Seyit Camtepe 等CVPR 2024
- Enhancing Adversarial Training with Second-Order Statistics of WeightsGaojie Jin, Xinping Yi, Wei Huang, Sven Schewe 等CVPR 2022 · 被引用 46 次
