Probabilistic Margins for Instance Reweighting in Adversarial Training
Qizhou Wang, Feng Liu, Bo Han, Tongliang Liu, Chen Gong, Gang Niu, Mingyuan Zhou, Masashi Sugiyama
Abstract
Reweighting adversarial data during training has been recently shown to improve adversarial robustness, where data closer to the current decision boundaries are regarded as more critical and given larger weights. However, existing methods measuring the closeness are not very reliable: they are discrete and can take only a few values, and they are path-dependent, i.e., they may change given the same start and end points with different attack paths. In this paper, we propose three types of probabilistic margin (PM), which are continuous and path-independent, for measuring the aforementioned closeness and reweighting adversarial data. Specifically, a PM is defined as the difference between two estimated class-posterior probabilities, e.g., such the probability of the true label minus the probability of the most confusing label given some natural data. Though different PMs capture different geometric properties, all three PMs share a negative correlation with the vulnerability of data: data with larger/smaller PMs are safer/riskier and should have smaller/larger weights. Experiments demonstrate that PMs are reliable measurements and PM-based reweighting methods outperform state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ccd77c8-e07d-4429-9b5f-c52f48618219Cited by top-tier papers30
- LAS-AT: Adversarial Training with Learnable Attack StrategyXiaojun Jia, Yong Zhang, Baoyuan Wu, Ke Ma et al.CVPR 2022 · 140 citations
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial TrainingLue Tao, Lei Feng, Jinfeng Yi, Sheng-Jun Huang et al.NeurIPS 2021 · 90 citations
- Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?Zhanke Zhou, Rong Tao, Jianing Zhu, Yiwen Luo et al.NeurIPS 2024 · 74 citations
- Learning to Augment Distributions for Out-of-distribution DetectionQizhou Wang, Zhen Fang, Yonggang Zhang, Feng Liu et al.NeurIPS 2023 · 59 citations
- Watermarking for Out-of-distribution DetectionQizhou Wang, Feng Liu, Yonggang Zhang, Jing Zhang et al.NeurIPS 2022 · 44 citations
Builds on18
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
Related papers
- Geometry-aware Instance-reweighted Adversarial TrainingJingfeng Zhang, Jianing Zhu, Gang Niu, Bo Han et al.ICLR 2021 · 316 citations
- Towards Million-Scale Adversarial Robustness Evaluation With Stronger Individual AttacksYong Xie, Weijie Zheng, Hanxun Huang, Guangnan Ye et al.CVPR 2025
- Fast and Reliable Evaluation of Adversarial Robustness with Minimum-Margin AttackRuize Gao, Jiongxiao Wang, Kaiwen Zhou, Feng Liu et al.ICML 2022 · 22 citations
- Exploring and Exploiting Decision Boundary Dynamics for Adversarial RobustnessYuancheng Xu, Yanchao Sun, Micah Goldblum, Tom Goldstein et al.ICLR 2023 · 10 citations
- One-vs-the-Rest Loss to Focus on Important Samples in Adversarial TrainingSekitoshi Kanai, Shin'ya Yamaguchi, Masanori Yamada, Hiroshi Takahashi et al.ICML 2023 · 14 citations
