The Effectiveness of Random Forgetting for Robust Generalization
Vijaya Raghavan T. Ramkumar, Bahram Zonooz, Elahe Arani
摘要
Deep neural networks are susceptible to adversarial attacks, which can compromise their performance and accuracy. Adversarial Training (AT) has emerged as a popular approach for protecting neural networks against such attacks. However, a key challenge of AT is robust overfitting, where the network's robust performance on test data deteriorates with further training, thus hindering generalization. Motivated by the concept of active forgetting in the brain, we introduce a novel learning paradigm called "Forget to Mitigate Overfitting (FOMO)". FOMO alternates between the forgetting phase, which randomly forgets a subset of weights and regulates the model's information through weight reinitialization, and the relearning phase, which emphasizes learning generalizable features. Our experiments on benchmark datasets and adversarial attacks show that FOMO alleviates robust overfitting by significantly reducing the gap between the best and last robust test accuracy while improving the state-of-the-art robustness. Furthermore, FOMO provides a better trade-off between standard and robust accuracy, outperforming baseline adversarial methods. Finally, our framework is robust to AutoAttacks and increases generalization in many real-world scenarios. 1 * Contributed equally. 1 Code is available at https://github.com/NeurAI-Lab/FOMO . 2 Double descent is a phenomenon in deep learning where a model's test error initially increases, decreases, and then increases again as model complexity or dataset size increases (Nakkiran et al., 2021) .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
相关 Paper
- Exploring The Forgetting in Adversarial Training: A Novel Method for Enhancing RobustnessXianglu Wang, Hu DingICLR 2025
- Consistency Regularization for Adversarial RobustnessJihoon Tack, Sihyun Yu, Jongheon Jeong, Minseon Kim 等AAAI 2022 · 被引用 75 次
- Exploring Memorization in Adversarial TrainingYinpeng Dong, Ke Xu, Xiao Yang, Tianyu Pang 等ICLR 2022 · 被引用 84 次
- Advancing Example Exploitation Can Alleviate Critical Challenges in Adversarial TrainingYao Ge, Yun Li, Keji Han, Junyi Zhu 等ICCV 2023 · 被引用 6 次
- On the Over-Memorization During Natural, Robust and Catastrophic OverfittingRunqi Lin, Chaojian Yu, Bo Han, Tongliang LiuICLR 2024 · 被引用 21 次
