The Silent Manipulator: A Practical and Inaudible Backdoor Attack against Speech Recognition Systems
Zhicong Zheng, Xinfeng Li, Chen Yan, Xiaoyu Ji, Wenyuan Xu
Abstract
Backdoor Attacks have been shown to pose significant threats to automatic speech recognition systems (ASRs). Existing success largely assumes backdoor triggering in the digital domain, or the victim will not notice the presence of triggering sounds in the physical domain. However, in practical victim-present scenarios, the over-the-air distortion of the backdoor trigger and the victim awareness raised by its audibility may invalidate such attacks. In this paper, we propose SMA, an inaudible grey-box backdoor attack that can be generalized to real-world scenarios where victims are present by exploiting both the vulnerability of microphones and neural networks. Specifically, we utilize the nonlinear effects of microphones to inject an inaudible ultrasonic trigger. To accurately characterize the microphone response to the crafted ultrasound, we construct a novel nonlinear transfer function for effective optimization. We also design optimization objectives to ensure triggers' robustness in the physical world and transferability on unseen ASR models. In practice, SMA can bypass the microphone's built-in filters and human perception, activating the implanted trigger in the ASRs inaudibly, regardless of whether the user is speaking. Extensive experiments show that the attack success rate of SMA can reach nearly 100% in the digital domain and over 85% against most microphones in the physical domains by only poisoning about 0.5% of the training audio dataset. Moreover, our attack can resist typical defense countermeasures to backdoor attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- SafeEar: Content Privacy-Preserving Audio Deepfake DetectionXinfeng Li, Kai Li, Yifan Zheng, Chen Yan et al.CCS 2024 · 26 citations
- AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language ModelsKai Li, Can Shen, Yile Liu, Jirui Han et al.ICLR 2026 · 17 citations
- SoK: Understanding the Fundamentals and Implications of Sensor Out-of-band VulnerabilitiesShilin Xiao, Wenjun Zhu, Yan Jiang, Kai Wang et al.NDSS 2026 · 3 citations
- Speed Master: Quick or Slow Play to Attack Speaker RecognitionZhe Ye, Wenjie Zhang, Ying Ren, Xiangui Kang et al.AAAI 2025 · 1 citation
- PhyFuzz: Detecting Sensor Vulnerabilities with Physical Signal FuzzingZhicong Zheng, Jinghui Wu, Shilin Xiao, Yanze Ren et al.NDSS 2026
Builds on7
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- DolphinAttack: Inaudible Voice CommandsGuoming Zhang, Chen Yan, Xiaoyu Ji, Tianchen Zhang et al.CCS 2017 · 753 citations
- TrojDRL: Evaluation of Backdoor Attacks on Deep Reinforcement LearningPanagiota Kiourti, Kacper Wardega, Susmit Jha, Wenchao LiDAC 2020 · 72 citations
- AdvDoor: adversarial backdoor attack of deep learning systemQuan Zhang, Yifeng Ding, Yongqiang Tian, Jianmin Guo et al.ISSTA 2021 · 57 citations
Related papers
- Opportunistic Backdoor Attacks: Exploring Human-imperceptible Vulnerabilities on Speech Recognition SystemsQiang Liu, Tongqing Zhou, Zhiping Cai, Yonghao TangACM MM 2022 · 33 citations
- FlowMur: A Stealthy and Practical Audio Backdoor Attack with Limited KnowledgeJiahe Lan, Jie Wang, Baochen Yan, Zheng Yan et al.S&P 2024 · 23 citations
- Devil in the Room: Triggering Audio Backdoors in the Physical WorldMeng Chen, Xiangyu Xu, Li Lu, Zhongjie Ba et al.USENIX Security 2024 · 7 citations
- Modulation-Based Backdoors: Leveraging Amplitude and Frequency Patterns to Attack Speaker RecognitionHanbo Cai, Pengcheng Zhang, Yan Xiao, De Li et al.AAAI 2026
- Inaudible Backdoor Attack via Stealthy Frequency Trigger Injection in Audio SpectrogramTianfang Zhang, Huy Phan, Zijie Tang, Cong Shi et al.MobiCom 2024 · 8 citations
