Adversarial Training for Defense Against Label Poisoning Attacks
Melis Ilayda Bal, Volkan Cevher, Michael Muehlebach
Abstract
As machine learning models grow in complexity and increasingly rely on publicly sourced data, such as the human-annotated labels used in training large language models, they become more vulnerable to label poisoning attacks. These attacks, in which adversaries subtly alter the labels within a training dataset, can severely degrade model performance, posing significant risks in critical applications. In this paper, we propose FLORAL, a novel adversarial training defense strategy based on support vector machines (SVMs) to counter these threats. Utilizing a bilevel optimization framework, we cast the training process as a non-zero-sum Stackelberg game between an attacker, who strategically poisons critical training labels, and the model, which seeks to recover from such attacks. Our approach accommodates various model architectures and employs a projected gradient descent algorithm with kernel SVMs for adversarial training. We provide a theoretical analysis of our algorithm's convergence properties and empirically evaluate FLORAL's effectiveness across diverse classification tasks. Compared to robust baselines and foundation models such as RoBERTa, FLORAL consistently achieves higher robust accuracy under increasing attacker budgets. These results underscore the potential of FLORAL to enhance the resilience of machine learning models against label poisoning threats, thereby ensuring robust classification in adversarial settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ac2ea5c-b803-489a-9eaa-dcbe87d1307bCited by top-tier papers2
- Safety-Efficacy Trade Off: Robustness against Data-PoisoningDiego Granziol, Ulugbek AbdimanabovICML 2026 · 1 citation
- On the Fragility of Data Attribution When Learning Is DistributedXian Gao, Bo Hui, MIN-TE SUN, Wei-Shinn KuICML 2026
Builds on13
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- A Closer Look at Accuracy vs. RobustnessYao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Ruslan Salakhutdinov et al.NeurIPS 2020 · 336 citations
- Geometry-aware Instance-reweighted Adversarial TrainingJingfeng Zhang, Jianing Zhu, Gang Niu, Bo Han et al.ICLR 2021 · 316 citations
- Adversarial Examples Make Strong PoisonsLiam Fowl, Micah Goldblum, Ping-yeh Chiang, Jonas Geiping et al.NeurIPS 2021 · 185 citations
- Certified Robustness to Label-Flipping Attacks via Randomized SmoothingElan Rosenfeld, Ezra Winston, Pradeep Ravikumar, J. Zico KolterICML 2020 · 182 citations
Related papers
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu et al.S&P 2018 · 867 citations
- Data Poisoning Attacks Against Multimodal EncodersZiqing Yang, Xinlei He, Zheng Li, Michael Backes et al.ICML 2023 · 74 citations
- Theory of Continual Learning Against Data Poisoning AttacksYiting Hu, Lingjie DuanICML 2026
- PoisonedEncoder: Poisoning the Unlabeled Pre-training Data in Contrastive LearningHongbin Liu, Jinyuan Jia, Neil Zhenqiang GongUSENIX Security 2022
- AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference OptimizationChaohu Liu, Tianyi Gui, Yu Liu, Linli XuICLR 2026 · 9 citations
