Robust Overfitting Does Matter: Test-Time Adversarial Purification with FGSM
Linyu Tang, Lei Zhang
Abstract
Numerous studies have demonstrated the susceptibility of deep neural networks (DNNs) to subtle adversarial perturbations, prompting the development of many advanced adversarial defense methods aimed at mitigating adversarial attacks. Current defense strategies usually train DNNs for a specific adversarial attack method and can achieve good robustness in defense against this type of adversarial attack. Nevertheless, when subjected to evaluations involving unfamiliar attack modalities, empirical evidence reveals a pronounced deterioration in the robustness of DNNs. Meanwhile, there is a trade-off between the classification accuracy of clean examples and adversarial examples. Most defense methods often sacrifice the accuracy of clean examples in order to improve the adversarial robustness of DNNs. To alleviate these problems and enhance the overall robust generalization of DNNs, we propose the Test-Time Pixel-Level Adversarial Purification (TPAP) method. This approach is based on the robust overfitting characteristic of DNNs to the fast gradient sign method (FGSM) on training and test datasets. It utilizes FGSM for adversarial purification, to process images for purifying unknown adversarial perturbations from pixels at testing time in a "counter changes with changelessness" manner, thereby enhancing the defense capability of DNNs against various unknown adversarial attacks. Extensive experimental results show that our method can effectively improve both overall robust generalization of DNNs, notably over previous methods. Code is available https://github. com/tly18/TPAP . 0 25 50 75 100 Epoch 0.4 0.6 0.8 1.0 Accuracy Training…results…on…clean…examples FGSM-AT-x train c CW-AT-x train c PGD-AT-x train c STA-AT-x train c (a) 0 25 50 75 100 Epoch 0.25 0.50 0.75 1.00 Accuracy Training…on…adversarial…examples FGSM-AT-x train FGSM CW-AT-x train CW2 PGD-AT-x train PGD STA-AT-x train STA (b) 0 25 50 75 100 Epoch 0.4 0.6 0.8 Accuracy Test…results…on…clean…examples FGSM-AT-x test c CW-AT-x test c PGD-AT-x test c STA-AT-x test c (c) 0 25 50 75 100 Epoch 0.0 0.5
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e401ef4-ff2a-4746-b231-fee91f4187b6Cited by top-tier papers4
- One Surrogate to Fool Them All: Universal, Transferable, and Targeted Adversarial Attacks with CLIPBinyan Xu, Xilin Dai, Di Tang, Kehuan ZhangCCS 2025 · 1 citation
- Divide and Conquer: Heterogeneous Noise Integration for Diffusion-based Adversarial PurificationGaozheng Pei, Shaojie Lyu, Gong Chen, Ke Ma et al.CVPR 2025
- Diffusion-based Adversarial Purification from the Perspective of the Frequency DomainGaozheng Pei, Ke Ma, Yingfei Sun, Qianqian Xu et al.ICML 2025
- DRFGD: Disentangled Representation-Focused Generative Defense for Attack-Tolerant Cross-Modal HashingZhongqing Yu, Xin Liu, Yiu-ming Cheung, Zhikai Hu et al.AAAI 2026
Builds on22
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
Related papers
- Adversarial Training on Purification (AToP): Advancing Both Robustness and GeneralizationGuang Lin, Chao Li, Jianhai Zhang, Toshihisa Tanaka et al.ICLR 2024 · 25 citations
- Online Adversarial Purification based on Self-supervised LearningChanghao Shi, Chester Holtz, Gal MishneICLR 2021 · 63 citations
- Feature Purification: How Adversarial Training Performs Robust Deep LearningZeyuan Allen-Zhu, Yuanzhi LiFOCS 2021 · 83 citations
- Adversary Aware Optimization for Robust DefenseDaniel Wesego, Pedram RooshenasNeurIPS 2025 · 3 citations
- Hilbert-Based Generative Defense for Adversarial ExamplesYang Bai, Yan Feng, Yisen Wang, Tao Dai et al.ICCV 2019 · 62 citations
