Mithridates: Auditing and Boosting Backdoor Resistance of Machine Learning Pipelines
Eugene Bagdasarian, Vitaly Shmatikov
Abstract
Machine learning (ML) models trained on data from potentially untrusted sources are vulnerable to poisoning. A small, maliciously crafted subset of the training inputs can cause the model to learn a "backdoor" task (e.g., misclassify inputs with a certain feature) in addition to its main task. Recent research proposed many hypothetical backdoor attacks whose efficacy heavily depends on the configuration and training hyperparameters of the target model. At the same time, state-of-the-art defenses require massive changes to the existing ML pipelines and may protect only against some attacks. Given the variety of potential backdoor attacks, ML engineers who are not security experts have no way to measure how vulnerable their current training pipelines are, nor do they have a practical way to compare training configurations so as to pick the more resistant ones. Deploying a defense may not be a realistic option, either. It requires evaluating and choosing from among dozens of research papers, completely re-engineering the pipeline as required by the chosen defense, and then repeating the process if the defense disrupts normal model training (while providing theoretical protection against an unknown subset of hypothetical threats). In this paper, we aim to provide ML engineers with pragmatic tools to audit the backdoor resistance of their training pipelines and to compare different training configurations, to help choose one that best balances accuracy and security. First, we propose a universal, attack-agnostic resistance metric based on the minimum number of training inputs that must be compromised before the model learns any backdoor. Second, we design, implement, and evaluate Mithridates, 1 a multi-stage approach that integrates backdoor resistance into the training-configuration search. ML developers already rely on hyperparameter search to find configurations that maximize the model's accuracy. Mithridates extends this standard tool to balance accuracy and resistance without disruptive changes to the training pipeline. We show that hyperparameters found by Mithridates increase resistance to multiple types of backdoor attacks by 3-5x with only a slight impact on accuracy. We also discuss extensions to AutoML and federated learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7961ed4f-9dcf-4570-bfbd-5191402489a1Cited by top-tier papers2
- Tuning for Two Adversaries: Enhancing the Robustness Against Transfer and Query-Based Attacks Using Hyperparameter TuningPascal Zimmer, Ghassan KarameAAAI 2026
- On Hyperparameters and Backdoor-Resistance in Horizontal Federated LearningSimon Lachnit, Ghassan KarameCCS 2025
Builds on50
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Attack of the Tails: Yes, You Really Can Backdoor Federated LearningHongyi Wang, Kartik Sreenivasan, Shashank Rajput, Harit Vishwakarma et al.NeurIPS 2020 · 862 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
Related papers
- Towards A Proactive ML Approach for Detecting Backdoor Poison SamplesXiangyu Qi, Tinghao Xie, Jiachen T. Wang, Tong Wu et al.USENIX Security 2023
- Mitigating Backdoor Attack by Injecting Proactive Defensive BackdoorShaokui Wei, Hongyuan Zha, Baoyuan WuNeurIPS 2024 · 20 citations
- Exploring the Orthogonality and Linearity of Backdoor AttacksKaiyuan Zhang, Siyuan Cheng, Guangyu Shen, Guanhong Tao et al.S&P 2024 · 11 citations
- BagFlip: A Certified Defense Against Data PoisoningYuhao Zhang, Aws Albarghouthi, Loris D'AntoniNeurIPS 2022 · 32 citations
- PoisonSpot: Precise Spotting of Clean-Label Backdoors via Fine-Grained Training Provenance TrackingPhilemon Hailemariam, Birhanu EsheteCCS 2025
