Focus on Hiders: Exploring Hidden Threats for Enhancing Adversarial Training
Qian Li, Yuxiao Hu, Yinpeng Dong, Dongxiao Zhang, Yuntian Chen
Abstract
Adversarial training is often formulated as a min-max problem, however, concentrating only on the worst adversarial examples causes alternating repetitive confusion of the model, i.e., previously defended or correctly classified samples are not defensible or accurately classifiable in subsequent adversarial training. We characterize such nonignorable samples as "hiders", which reveal the hidden high-risk regions within the secure area obtained through adversarial training and prevent the model from finding the real worst cases. We demand the model to prevent hiders when defending against adversarial examples for improving accuracy and robustness simultaneously. By rethinking and redefining the min-max optimization problem for adversarial training, we propose a generalized adversarial training algorithm called Hider-Focused Adversarial Training (HFAT). HFAT introduces the iterative evolution optimization strategy to simplify the optimization problem and employs an auxiliary model to reveal hiders, effectively combining the optimization directions of standard adversarial training and prevention hiders. Furthermore, we introduce an adaptive weighting mechanism that facilitates the model in adaptively adjusting its focus between adversarial examples and hiders during different training periods. We demonstrate the effectiveness of our method based on extensive experiments, and ensure that HFAT can provide higher robustness and accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5aa60b32-86a0-43c3-b2e1-7644e0820c5dCited by top-tier papers4
- GenSR: Symbolic regression based on equation generative spaceQian Li, Yuxiao Hu, Juncheng Liu, Yuntian ChenICLR 2026 · 7 citations
- Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty QuantificationTao Huang, Rui Wang, Xiaofei Liu, Yi Qin et al.ICLR 2026 · 4 citations
- Rethinking Invariance Regularization in Adversarial Training to Improve Robustness-Accuracy Trade-offFuta Kai Waseda, Ching-Chun Chang, Isao EchizenICLR 2025
- Weakly Supervised Contrastive Adversarial Training for Learning Robust Features from Semi-supervised DataLilin Zhang, Chengpei Wu, Ning YangCVPR 2025
Builds on13
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
Related papers
- Attacks Which Do Not Kill Training Make Adversarial Learning StrongerJingfeng Zhang, Xilie Xu, Bo Han, Gang Niu et al.ICML 2020 · 452 citations
- Reducing Excessive Margin to Achieve a Better Accuracy vs. Robustness Trade-offRahul Rade, Seyed-Mohsen Moosavi-DezfooliICLR 2022 · 166 citations
- Are Adversarial Examples Created Equal? A Learnable Weighted Minimax Risk for Robustness under Non-uniform AttacksHuimin Zeng, Chen Zhu, Tom Goldstein, Furong HuangAAAI 2021 · 21 citations
- Robust Local Features for Improving the Generalization of Adversarial TrainingChuanbiao Song, Kun He, Jiadong Lin, Liwei Wang et al.ICLR 2020 · 78 citations
- Understanding and Improving Ensemble Adversarial DefenseYian Deng, Tingting MuNeurIPS 2023 · 37 citations
