Enhancing Adversarial Defense by k-Winners-Take-All
Chang Xiao, Peilin Zhong, Changxi Zheng
摘要
We propose a simple change to existing neural network structures for better defending against gradient-based adversarial attacks. Instead of using popular activation functions (such as ReLU), we advocate the use of k-Winners-Take-All (k-WTA) activation, a C0 discontinuous function that purposely invalidates the neural network model's gradient at densely distributed input data points. The proposed k-WTA activation can be readily used in nearly all existing networks and training methods with no significant overhead. Our proposal is theoretically rationalized. We analyze why the discontinuities in k-WTA networks can largely prevent gradient-based search of adversarial examples and why they at the same time remain innocuous to the network training. This understanding is also empirically backed. We test k-WTA activation on various network structures optimized by a training method, be it adversarial training or not. In all cases, the robustness of k-WTA networks outperforms that of traditional networks under white-box attacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Adversarial Purification with Score-based Generative ModelsJongmin Yoon, Sung Ju Hwang, Juho LeeICML 2021 · 被引用 200 次
- Learnable Boundary Guided Adversarial TrainingJiequan Cui, Shu Liu, Liwei Wang, Jiaya JiaICCV 2021 · 被引用 152 次
- Meta Gradient Adversarial AttackZheng Yuan, Jie Zhang, Yunpei Jia, Chuanqi Tan 等ICCV 2021 · 被引用 95 次
- Mind the Box: l1-APGD for Sparse Adversarial Attacks on Image ClassifiersFrancesco Croce, Matthias HeinICML 2021 · 被引用 68 次
它引用的顶会 Paper2
相关 Paper
- Discrete Adversarial Attack to Models of CodeFengjuan Gao, Yu Wang, Ke WangPLDI 2023 · 被引用 23 次
- Low Curvature Activations Reduce Overfitting in Adversarial TrainingVasu Singla, Sahil Singla, Soheil Feizi, David JacobsICCV 2021 · 被引用 49 次
- Detecting Adversarial Samples Using Influence Functions and Nearest NeighborsGilad Cohen, Guillermo Sapiro, Raja GiryesCVPR 2020
- One Man's Trash Is Another Man's Treasure: Resisting Adversarial Examples by Adversarial ExamplesChang Xiao, Changxi ZhengCVPR 2020
- TRNAS: A Training-Free Robust Neural Architecture SearchYeming Yang, Qingling Zhu, Jianping Luo, Ka-Chun Wong 等ICCV 2025 · 被引用 1 次
