Audio-domain position-independent backdoor attack via unnoticeable triggers
Cong Shi, Tianfang Zhang, Zhuohang Li, Huy Phan, Tianming Zhao, Yan Wang, Jian Liu, Bo Yuan, Yingying Chen
Abstract
Deep learning models have become key enablers of voice user interfaces. With the growing trend of adopting outsourced training of these models, backdoor attacks, stealthy yet effective training-phase attacks, have gained increasing attention. They inject hidden trigger patterns through training set poisoning and overwrite the model's predictions in the inference phase. Research in backdoor attacks has been focusing on image classification tasks, while there have been few studies in the audio domain. In this work, we explore the severity of audio-domain backdoor attacks and demonstrate their feasibility under practical scenarios of voice user interfaces, where an adversary injects (plays) an unnoticeable audio trigger into live speech to launch the attack. To realize such attacks, we consider jointly optimizing the audio trigger and the target model in the training phase, deriving a position-independent, unnoticeable, and robust audio trigger. We design new data poisoning techniques and penalty-based algorithms that inject the trigger into randomly generated temporal positions in the audio input during training, rendering the trigger resilient to any temporal position variations. We further design an environmental sound mimicking technique to make the trigger resemble unnoticeable situational sounds and simulate played over-the-air distortions to improve the trigger's robustness during the joint optimization process. Extensive experiments on two important applications (i.e., speech command recognition and speaker recognition) demonstrate that our attack can achieve an average success rate of over 99% under both digital and physical attack settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- FlowMur: A Stealthy and Practical Audio Backdoor Attack with Limited KnowledgeJiahe Lan, Jie Wang, Baochen Yan, Zheng Yan et al.S&P 2024 · 23 citations
- MASTERKEY: Practical Backdoor Attack Against Speaker Verification SystemsHanqing Guo, Xun Chen, Junfeng Guo, Li Xiao et al.MobiCom 2023 · 14 citations
- CSTAR: Towards Compact and Structured Deep Neural Networks with Adversarial RobustnessHuy Phan, Miao Yin, Yang Sui, Bo Yuan et al.AAAI 2023 · 10 citations
- Distributed Backdoor Attacks on Federated Graph Learning and Certified DefensesYuxin Yang, Qiang Li, Jinyuan Jia, Yuan Hong et al.CCS 2024 · 8 citations
- Inaudible Backdoor Attack via Stealthy Frequency Trigger Injection in Audio SpectrogramTianfang Zhang, Huy Phan, Zijie Tang, Cong Shi et al.MobiCom 2024 · 8 citations
Builds on9
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 601 citations
- Who is Real Bob? Adversarial Attacks on Speaker Recognition SystemsGuangke Chen, Sen Chen, Lingling Fan, Xiaoning Du et al.S&P 2021 · 239 citations
Related papers
- Modulation-Based Backdoors: Leveraging Amplitude and Frequency Patterns to Attack Speaker RecognitionHanbo Cai, Pengcheng Zhang, Yan Xiao, De Li et al.AAAI 2026
- Devil in the Room: Triggering Audio Backdoors in the Physical WorldMeng Chen, Xiangyu Xu, Li Lu, Zhongjie Ba et al.USENIX Security 2024 · 7 citations
- Opportunistic Backdoor Attacks: Exploring Human-imperceptible Vulnerabilities on Speech Recognition SystemsQiang Liu, Tongqing Zhou, Zhiping Cai, Yonghao TangACM MM 2022 · 33 citations
- DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation ConstraintsZhendong Zhao, Xiaojun Chen, Yuexin Xuan, Ye Dong et al.CVPR 2022 · 72 citations
- Speed Master: Quick or Slow Play to Attack Speaker RecognitionZhe Ye, Wenjie Zhang, Ying Ren, Xiangui Kang et al.AAAI 2025 · 1 citation
