Adversarial Training on Purification (AToP): Advancing Both Robustness and Generalization
Guang Lin, Chao Li, Jianhai Zhang, Toshihisa Tanaka, Qibin Zhao
Abstract
The deep neural networks are known to be vulnerable to well-designed adversarial attacks. The most successful defense technique based on adversarial training (AT) can achieve optimal robustness against particular attacks but cannot generalize well to unseen attacks. Another effective defense technique based on adversarial purification (AP) can enhance generalization but cannot achieve optimal robustness. Meanwhile, both methods share one common limitation on the degraded standard accuracy. To mitigate these issues, we propose a novel pipeline to acquire the robust purifier model, named Adversarial Training on Purification (AToP), which comprises two components: perturbation destruction by random transforms (RT) and purifier model fine-tuned (FT) by adversarial loss. RT is essential to avoid overlearning to known attacks, resulting in the robustness generalization to unseen attacks, and FT is essential for the improvement of robustness. To evaluate our method in an efficient and scalable way, we conduct extensive experiments on CIFAR-10, CIFAR-100, and ImageNette to demonstrate that our method achieves optimal robustness and exhibits generalization ability against unseen attacks. Our code is available at https://github.com/glin2022/atop .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6ace370-ce54-4a94-aac7-7382d869bc08Cited by top-tier papers5
- VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language ModelsHefei Mei, Zirui Wang, Shen You, Minjing Dong et al.ICLR 2026 · 9 citations
- Divide and Conquer: Heterogeneous Noise Integration for Diffusion-based Adversarial PurificationGaozheng Pei, Shaojie Lyu, Gong Chen, Ke Ma et al.CVPR 2025
- One Stone, Two Birds: Enhancing Adversarial Defense Through the Lens of Distributional DiscrepancyJiacheng Zhang, Benjamin I. P. Rubinstein, Jingfeng Zhang, Feng LiuICML 2025
- Why Adversarially Train Diffusion Models?Maria Rosaria Briglia, Mujtaba Hussain Mirza, Giuseppe Lisanti, Iacopo MasiICLR 2026
- Diffusion-based Adversarial Purification from the Perspective of the Frequency DomainGaozheng Pei, Ke Ma, Yingfei Sun, Qianqian Xu et al.ICML 2025
Builds on17
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao et al.ICML 2022 · 663 citations
- Improving Robustness using Generated DataSven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg et al.NeurIPS 2021 · 384 citations
Related papers
- Robust Overfitting Does Matter: Test-Time Adversarial Purification with FGSMLinyu Tang, Lei ZhangCVPR 2024 · 13 citations
- MimicDiffusion: Purifying Adversarial Perturbation via Mimicking Clean Diffusion ModelKaiyu Song, Hanjiang Lai, Yan Pan, Jian YinCVPR 2024 · 11 citations
- Adversary Aware Optimization for Robust DefenseDaniel Wesego, Pedram RooshenasNeurIPS 2025 · 3 citations
- Feature Purification: How Adversarial Training Performs Robust Deep LearningZeyuan Allen-Zhu, Yuanzhi LiFOCS 2021 · 83 citations
- Self-supervised Adversarial Purification for Graph Neural NetworksWoohyun Lee, Hogun ParkICML 2025
