Try to Poison My Deep Learning Data? Nowhere to Hide Your Trajectory Spectrum!
Yansong Gao, Huaibing Peng, Hua Ma, Zhi Zhang, Shuo Wang, Rayne Holland, Anmin Fu, Minhui Xue, Derek Abbott
Abstract
—In the Data as a Service (DaaS) model, data cura-tors, such as commercial providers like Amazon Mechanical Turk, Appen, and TELUS International, aggregate quality data from numerous contributors and monetize it for deep learning (DL) model providers. However, malicious contributors can poison this data, embedding backdoors in the trained DL models. Existing methods for detecting poisoned samples face significant limitations: they often rely on reserved clean data; they are sensitive to the poisoning rate, trigger type, and backdoor type; and they are specific to classification tasks. These limitations hinder their practical adoption by data curators. This work, for the first time, investigates the training trajectory of poisoned samples in the spectrum domain , revealing distinctions from benign samples that are not apparent in the original non-spectrum domain. Building on this novel perspective, we propose Telltale to detect and sanitize poisoned samples as a one-time effort, addressing all of the aforementioned limitations of prior work. Through extensive experiments, Telltale demonstrates the ability to defeat both universal and challenging partial backdoor types without relying on any reserved clean data. Telltale is also validated to be agnostic to various trigger types, including the advanced clean-label trigger attack, Narcissus (CCS’2023). Moreover, Telltale proves effective across diverse data modalities (e.g., image, audio and text) and non-classification tasks (e.g., regression)—making it the only known training phase poisoned sample detection method applicable to non-classification tasks. In all our evaluations, Telltale achieves a detection accuracy (i.e., accurately identifying poisoned samples) of at least 95.52% and a false positive rate (i.e., falsely recognizing benign samples as poisoned ones) no higher than 0.61%. Comparisons with state-of-the-art methods, ASSET (Usenix’2023) and CT (Usenix’2023), further affirm Telltale ’s superior performance. More specifically, ASSET fails to handle partial backdoor types and incurs an unbearable false positive rate with clean/benign datasets common in practice, while CT fails against the Narcissus trigger. In contrast, Telltale proves highly effective across testing scenarios where prior work fails. The source code is released at https://github.com/MPaloze/Telltale.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on35
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- DBA: Distributed Backdoor Attacks against Federated LearningChulin Xie, Keli Huang, Pin-Yu Chen, Bo LiICLR 2020 · 901 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
- Invisible Backdoor Attack with Sample-Specific TriggersYuezun Li, Yiming Li, Baoyuan Wu, Longkang Li et al.ICCV 2021 · 639 citations
Related papers
- Backdoor Defense via Test-Time Detecting and RepairingJiyang Guan, Jian Liang, Ran HeCVPR 2024
- PoisonSpot: Precise Spotting of Clean-Label Backdoors via Fine-Grained Training Provenance TrackingPhilemon Hailemariam, Birhanu EsheteCCS 2025
- DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation ConstraintsZhendong Zhao, Xiaojun Chen, Yuexin Xuan, Ye Dong et al.CVPR 2022 · 72 citations
- CLEAR: Clean-up Sample-Targeted Backdoor in Neural NetworksLiuwan Zhu, Rui Ning, Chunsheng Xin, Chonggang Wang et al.ICCV 2021 · 14 citations
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 19 citations
