PureGen: Universal Data Purification for Train-Time Poison Defense via Generative Model Dynamics
Omead Pooladzandi, Sunay Bhat, Jeffrey Jiang, Alexander Branch, Gregory J. Pottie
Abstract
Train-time data poisoning attacks threaten machine learning models by introducing adversarial examples during training, leading to misclassification. Current defense methods often reduce generalization performance, are attack-specific, and impose significant training overhead. To address this, we introduce a set of universal data purification methods using a stochastic transform, , realized via iterative Langevin dynamics of Energy-Based Models (EBMs), Denoising Diffusion Probabilistic Models (DDPMs), or both. These approaches purify poisoned data with minimal impact on classifier generalization. Our specially trained EBMs and DDPMs provide state-of-the-art defense against various attacks (including Narcissus, Bullseye Polytope, Gradient Matching) on CIFAR-10, Tiny-ImageNet, and CINIC-10, without needing attack or classifier-specific information. We discuss performance trade-offs and show that our methods remain highly effective even with poisoned or distributionally shifted generative model training data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 187ce43e-a259-4055-a97a-575b178d341cCited by top-tier papers1
Ask how each one uses itBuilds on23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao et al.ICML 2022 · 663 citations
Related papers
- Adversarial Purification with Score-based Generative ModelsJongmin Yoon, Sung Ju Hwang, Juho LeeICML 2021 · 200 citations
- Not All Poisons are Created Equal: Robust Training against Data PoisoningYu Yang, Tian Yu Liu, Baharan MirzasoleimanICML 2022 · 45 citations
- Stochastic Security: Adversarial Defense Using Long-Run Dynamics of Energy-Based ModelsMitch Hill, Jonathan Craig Mitchell, Song-Chun ZhuICLR 2021 · 93 citations
- Friendly Noise against Adversarial Noise: A Powerful Defense against Data Poisoning AttackTian Yu Liu, Yu Yang, Baharan MirzasoleimanNeurIPS 2022 · 39 citations
- Text Adversarial Purification as Defense against Adversarial AttacksLinyang Li, Demin Song, Xipeng QiuACL 2023 · 11 citations
