Eliminating Catastrophic Overfitting Via Abnormal Adversarial Examples Regularization
Runqi Lin, Chaojian Yu, Tongliang Liu
摘要
Single-step adversarial training (SSAT) has demonstrated the potential to achieve both efficiency and robustness. However, SSAT suffers from catastrophic overfitting (CO), a phenomenon that leads to a severely distorted classifier, making it vulnerable to multi-step adversarial attacks. In this work, we observe that some adversarial examples generated on the SSAT-trained network exhibit anomalous behaviour, that is, although these training samples are generated by the inner maximization process, their associated loss decreases instead, which we named abnormal adversarial examples (AAEs). Upon further analysis, we discover a close relationship between AAEs and classifier distortion, as both the number and outputs of AAEs undergo a significant variation with the onset of CO. Given this observation, we re-examine the SSAT process and uncover that before the occurrence of CO, the classifier already displayed a slight distortion, indicated by the presence of few AAEs. Furthermore, the classifier directly optimizing these AAEs will accelerate its distortion, and correspondingly, the variation of AAEs will sharply increase as a result. In such a vicious circle, the classifier rapidly becomes highly distorted and manifests as CO within a few iterations. These observations motivate us to eliminate CO by hindering the generation of AAEs. Specifically, we design a novel method, termed Abnormal Adversarial Examples Regularization (AAER), which explicitly regularizes the variation of AAEs to hinder the classifier from becoming distorted. Extensive experiments demonstrate that our method can effectively eliminate CO and further boost adversarial robustness with negligible additional computational overhead. Our implementation can be found at https://github.com/tmllab/2023_NeurIPS_AAER .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context LearningZhuo Huang, Chang Liu, Yinpeng Dong, Hang Su 等ICML 2024 · 被引用 31 次
- On the Over-Memorization During Natural, Robust and Catastrophic OverfittingRunqi Lin, Chaojian Yu, Bo Han, Tongliang LiuICLR 2024 · 被引用 21 次
- Layer-Aware Analysis of Catastrophic Overfitting: Revealing the Pseudo-Robust Shortcut DependencyRunqi Lin, Chaojian Yu, Bo Han, Hang Su 等ICML 2024 · 被引用 9 次
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?Junchi Yu, Yujie Liu, Jindong Gu, Philip H. S. Torr 等NeurIPS 2025 · 被引用 8 次
- Mobile-VTON: High-Fidelity On-Device Virtual Try-OnZhenchen Wan, Ce Chen, Runqi Lin, Jiaxin Huang 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper18
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 被引用 1,765 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- Understanding and Improving Fast Adversarial TrainingMaksym Andriushchenko, Nicolas FlammarionNeurIPS 2020 · 被引用 366 次
相关 Paper
- Understanding Catastrophic Overfitting in Single-step Adversarial TrainingHoki Kim, Woojin Lee, Jaewook LeeAAAI 2021 · 被引用 135 次
- Efficient local linearity regularization to overcome catastrophic overfittingElías Abad-Rocamora, Fanghui Liu, Grigorios Chrysos, Pablo M. Olmos 等ICLR 2024 · 被引用 9 次
- Taxonomy Driven Fast Adversarial TrainingKun Tong, Chengze Jiang, Jie Gui, Yuan CaoAAAI 2024 · 被引用 2 次
- Mitigating Catastrophic Overfitting in Fast Adversarial Training via Label Information EliminationChao Pan, Ke Tang, Qing Li, Xin YaoICCV 2025 · 被引用 1 次
- SORA: Free Second-Order Attacks in Fast Adversarial TrainingMazdak Teymourian, Ramtin Moslemi, Farzan Rahmani, Mohammad H RohbanICML 2026
