Distilling Cognitive Backdoor Patterns within an Image
Hanxun Huang, Xingjun Ma, Sarah Monazam Erfani, James Bailey
Abstract
This paper proposes a simple method to distill and detect backdoor patterns within an image: Cognitive Distillation (CD). The idea is to extract the "minimal essence" from an input image responsible for the model's prediction. CD optimizes an input mask to extract a small pattern from the input image that can lead to the same model output (i.e., logits or deep features). The extracted pattern can help understand the cognitive mechanism of a model on clean vs. backdoor images and is thus called a Cognitive Pattern (CP). Using CD and the distilled CPs, we uncover an interesting phenomenon of backdoor attacks: despite the various forms and sizes of trigger patterns used by different attacks, the CPs of backdoor samples are all surprisingly and suspiciously small. One thus can leverage the learned mask to detect and remove backdoor examples from poisoned training datasets. We conduct extensive experiments to show that CD can robustly detect a wide range of advanced backdoor attacks. We also show that CD can potentially be applied to help detect potential biases from face datasets. Code is available at https://github.com/HanxunH/CognitiveDistillation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0db7e9e4-84ab-4b67-acc1-40a2d611ce18Cited by top-tier papers18
- Towards Reliable and Efficient Backdoor Trigger Inversion via Decoupling Benign FeaturesXiong Xu, Kunzhe Huang, Yiming Li, Zhan Qin et al.ICLR 2024 · 59 citations
- IBD-PSC: Input-level Backdoor Detection via Parameter-oriented Scaling ConsistencyLinshan Hou, Ruili Feng, Zhongyun Hua, Wei Luo et al.ICML 2024 · 52 citations
- SampDetox: Black-box Backdoor Defense via Perturbation-based Sample DetoxificationYanxin Yang, Chentao Jia, Dengke Yan, Ming Hu et al.NeurIPS 2024 · 20 citations
- Breaking Barriers in Physical-World Adversarial Examples: Improving Robustness and Transferability via Robust FeatureYichen Wang, Yuxuan Chou, Ziqi Zhou, Hangtao Zhang et al.AAAI 2025 · 20 citations
- From Trojan Horses to Castle Walls: Unveiling Bilateral Data Poisoning Effects in Diffusion ModelsZhuoshi Pan, Yuguang Yao, Gaowen Liu, Bingquan Shen et al.NeurIPS 2024 · 15 citations
Builds on41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 1,765 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
Related papers
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.ICLR 2021 · 548 citations
- Rethinking Backdoor Attacks on Dataset Distillation: A Kernel Method PerspectiveMing-Yu Chung, Sheng-Yen Chou, Chia-Mu Yu, Pin-Yu Chen et al.ICLR 2024 · 12 citations
- BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge DistillationZhengxian Wu, Juan Wen, Wanli Peng, Yinghan Zhou et al.AAAI 2026 · 2 citations
- Pre-activation Distributions Expose Backdoor NeuronsRunkai Zheng, Rongjun Tang, Jianze Li, Li LiuNeurIPS 2022 · 51 citations
- Backdoor Attacks Against Dataset DistillationYugeng Liu, Zheng Li, Michael Backes, Yun Shen et al.NDSS 2023
