Adversarial Unlearning: Reducing Confidence Along Adversarial Directions
Amrith Setlur, Benjamin Eysenbach, Virginia Smith, Sergey Levine
摘要
Supervised learning methods trained with maximum likelihood objectives often overfit on training data. Most regularizers that prevent overfitting look to increase confidence on additional examples (e.g., data augmentation, adversarial training), or reduce it on training data (e.g., label smoothing). In this work we propose a complementary regularization strategy that reduces confidence on self-generated examples. The method, which we call RCAD (Reducing Confidence along Adversarial Directions), aims to reduce confidence on out-of-distribution examples lying along directions adversarially chosen to increase training loss. In contrast to adversarial training, RCAD does not try to robustify the model to output the original label, but rather regularizes it to have reduced confidence on points generated using much larger perturbations than in conventional adversarial training. RCAD can be easily integrated into training pipelines with a few lines of code. Despite its simplicity, we find on many classification benchmarks that RCAD can be added to existing techniques (e.g., label smoothing, MixUp training) to increase test accuracy by 1-3% in absolute value, with more significant gains in the low data regime. We also provide a theoretical analysis that helps to explain these benefits in simplified settings, showing that RCAD can provably help the model unlearn spurious features in the training data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Fast Federated Machine Unlearning with Nonlinear Functional TheoryTianshi Che, Yang Zhou, Zijie Zhang, Lingjuan Lyu 等ICML 2023 · 被引用 77 次
- Ignorance is Bliss: Robust Control via Information GatingManan Tomar, Riashat Islam, Matthew E. Taylor, Sergey Levine 等NeurIPS 2023 · 被引用 14 次
- Towards Safe Machine Unlearning: A Paradigm that Mitigates Performance DegradationShanshan Ye, Jie Lu, Guangquan ZhangWWW 2025 · 被引用 13 次
- Contrastive Novelty-Augmented Learning: Anticipating Outliers with Large Language ModelsAlbert Xu, Xiang Ren, Robin JiaACL 2023 · 被引用 1 次
- Tool Unlearning for Tool-Augmented LLMsJiali Cheng, Hadi AmiriICML 2025
它引用的顶会 Paper15
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Semi-Supervised Domain Adaptation via Minimax EntropyKuniaki Saito, Donghyun Kim, Stan Sclaroff, Trevor Darrell 等ICCV 2019 · 被引用 725 次
相关 Paper
- SmoothMix: Training Confidence-calibrated Smoothed Classifiers for Certified RobustnessJongheon Jeong, Sejun Park, Minkyu Kim, Heung-Chang Lee 等NeurIPS 2021 · 被引用 70 次
- MaxSup: Overcoming Representation Collapse in Label SmoothingYuxuan Zhou, Heng Li, Zhi-Qi Cheng, Xudong Yan 等NeurIPS 2025 · 被引用 5 次
- Boundary thickness and robustness in learning modelsYaoqing Yang, Rajiv Khanna, Yaodong Yu, Amir Gholami 等NeurIPS 2020 · 被引用 53 次
- MaxUp: Lightweight Adversarial Training With Data Augmentation Improves Neural Network TrainingChengyue Gong, Tongzheng Ren, Mao Ye, Qiang LiuCVPR 2021
- Towards Understanding the Regularization of Adversarial Robustness on Neural NetworksYuxin Wen, Shuai Li, Kui JiaICML 2020 · 被引用 25 次
