Seal Your Backdoor with Variational Defense
Ivan Sabolic, Matej Grcic, Sinisa Segvic
Abstract
We propose VIBE, a model-agnostic framework that trains classifiers resilient to backdoor attacks. The key concept behind our approach is to treat malicious inputs and corrupted labels from the training dataset as observed random variables, while the actual clean labels are latent. VIBE then recovers the corresponding latent clean label posterior through variational inference. The resulting training procedure follows the expectation-maximization (EM) algorithm. The E-step infers the clean pseudolabels by solving an entropy-regularized optimal transport problem, while the M-step updates the classifier parameters via gradient descent. Being modular, VIBE can seamlessly integrate with recent advancements in self-supervised representation learning, which enhance its ability to resist backdoor attacks. We experimentally validate the method effectiveness against contemporary backdoor attacks on standard datasets, a large-scale setup with 1k classes, and a dataset poisoned with multiple attacks. VIBE consistently outperforms previous defenses across all tested scenarios.
1 Some attacks do not alter the labels [88]. However, our experiments show that they are much easier to defend from.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e4945d2-14c8-4b1c-bc83-a37a3c606721Cited by top-tier papers2
- BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor AttacksIvan Sabolic, Marin Oršić, Josip Šarić, Sven LoncaricICML 2026
- From Internal Diagnosis to External Auditing: A VLM-Driven Paradigm for Data-Free Online Backdoor DefenseBinyan Xu, Fan YANG, Xilin Dai, Di Tang et al.ICML 2026
Builds on62
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
Related papers
- Backdoor Attacks on Self-Supervised LearningAniruddha Saha, Ajinkya Tejankar, Soroush Abbasi Koohpayegani, Hamed PirsiavashCVPR 2022 · 77 citations
- Distribution Preserving Backdoor Attack in Self-supervised LearningGuanhong Tao, Zhenting Wang, Shiwei Feng, Guangyu Shen et al.S&P 2024 · 32 citations
- Backdoor Defense via Decoupling the Training ProcessKunzhe Huang, Yiming Li, Baoyuan Wu, Zhan Qin et al.ICLR 2022 · 253 citations
- Revisiting the Assumption of Latent Separability for Backdoor DefensesXiangyu Qi, Tinghao Xie, Yiming Li, Saeed Mahloujifar et al.ICLR 2023
- An Embarrassingly Simple Backdoor Attack on Self-supervised LearningChangjiang Li, Ren Pang, Zhaohan Xi, Tianyu Du et al.ICCV 2023 · 54 citations
