Learning to Generate Noise for Multi-Attack Robustness
Divyam Madaan, Jinwoo Shin, Sung Ju Hwang
Abstract
Adversarial learning has emerged as one of the successful techniques to circumvent the susceptibility of existing methods against adversarial perturbations. However, the majority of existing defense methods are tailored to defend against a single category of adversarial perturbation (e.g. -attack). In safety-critical applications, this makes these methods extraneous as the attacker can adopt diverse adversaries to deceive the system. Moreover, training on multiple perturbations simultaneously significantly increases the computational overhead during training. To address these challenges, we propose a novel meta-learning framework that explicitly learns to generate noise to improve the model's robustness against multiple types of attacks. Its key component is Meta Noise Generator (MNG) that outputs optimal noise to stochastically perturb a given sample, such that it helps lower the error on diverse adversarial perturbations. By utilizing samples generated by MNG, we train a model by enforcing the label consistency across multiple perturbations. We validate the robustness of models trained by our scheme on various datasets and against a wide variety of perturbations, demonstrating that it significantly outperforms the baselines across multiple perturbations with a marginal computational cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 704785a8-a385-42c8-abb9-b43791ab738fCited by top-tier papers10
- Efficient Test-Time Model Adaptation without ForgettingShuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Yaofo Chen et al.ICML 2022 · 579 citations
- Robust Prompt Optimization for Defending Language Models Against Jailbreaking AttacksAndy Zhou, Bo Li, Haohan WangNeurIPS 2024 · 198 citations
- Mind the Box: l1-APGD for Sparse Adversarial Attacks on Image ClassifiersFrancesco Croce, Matthias HeinICML 2021 · 68 citations
- Friendly Noise against Adversarial Noise: A Powerful Defense against Data Poisoning AttackTian Yu Liu, Yu Yang, Baharan MirzasoleimanNeurIPS 2022 · 39 citations
- Adversarial Robustness against Multiple and Single lp-Threat Models via Quick Fine-Tuning of Robust ClassifiersFrancesco Croce, Matthias HeinICML 2022 · 26 citations
Builds on11
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
Related papers
- MultiRobustBench: Benchmarking Robustness Against Multiple AttacksSihui Dai, Saeed Mahloujifar, Chong Xiang, Vikash Sehwag et al.ICML 2023 · 11 citations
- Combining Adversaries with Anti-adversaries in TrainingXiaoling Zhou, Nan Yang, Ou WuAAAI 2023 · 12 citations
- Adversarial Attack Generation Empowered by Min-Max OptimizationJingkang Wang, Tianyun Zhang, Sijia Liu, Pin-Yu Chen et al.NeurIPS 2021 · 49 citations
- Adversarial Attacks and Robust Training for Hypergraph Neural NetworksNaheed Anjum Arafat, Debabrota Basu, Yulia Gel, Danda RawatICML 2026
- Adversarial Defense via Learning to Generate Diverse AttacksYunseok Jang, Tianchen Zhao, Seunghoon Hong, Honglak LeeICCV 2019 · 88 citations
