Hear "No Evil", See "Kenansville"*: Efficient and Transferable Black-Box Attacks on Speech Recognition and Voice Identification Systems
Hadi Abdullah, Muhammad Sajidur Rahman, Washington Garcia, Kevin Warren, Anurag Swarnim Yadav, Tom Shrimpton, Patrick Traynor
Abstract
Automatic speech recognition and voice identification systems are being deployed in a wide array of applications, from providing control mechanisms to devices lacking traditional interfaces, to the automatic transcription of conversations and authentication of users. Many of these applications have significant security and privacy considerations. We develop attacks that force mistranscription and misidentification in state of the art systems, with minimal impact on human comprehension. Processing pipelines for modern systems are comprised of signal preprocessing and feature extraction steps, whose output is fed to a machine-learned model. Prior work has focused on the models, using white-box knowledge to tailor model-specific attacks. We focus on the pipeline stages before the models, which (unlike the models) are quite similar across systems. As such, our attacks are black-box and transferable, and demonstrably achieve mistranscription and misidentification rates as high as 100% by modifying only a few frames of audio. We perform a study via Amazon Mechanical Turk demonstrating that there is no statistically significant difference between human perception of regular and perturbed audio. Our findings suggest that models may learn aspects of speech that are generally not perceived by human subjects, but that are crucial for model accuracy. We also find that certain English language phonemes (in particular, vowels) are significantly more susceptible to our attack. We show that the attacks are effective when mounted over cellular networks, where signals are subject to degradation due to transcoding, jitter, and packet loss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 839fdcde-e9ee-4a7a-a498-af26bb03bf84Cited by top-tier papers18
- Who is Real Bob? Adversarial Attacks on Speaker Recognition SystemsGuangke Chen, Sen Chen, Lingling Fan, Xiaoning Du et al.S&P 2021 · 239 citations
- VoiceBlock: Privacy through Real-Time Adversarial Attacks with Audio-to-Audio ModelsPatrick O'Reilly, Andreas Bugler, Keshav Bhandari, Max Morrison et al.NeurIPS 2022 · 18 citations
- When Evil Calls: Targeted Adversarial Voice over IP NetworkHan Liu, Zhiyuan Yu, Mingming Zha, XiaoFeng Wang et al.CCS 2022 · 13 citations
- Perception-Aware Attack: Creating Adversarial Music via Reverse-Engineering Human PerceptionRui Duan, Zhe Qu, Shangqing Zhao, Leah Ding et al.CCS 2022 · 8 citations
- Defending against Adversarial Audio via Diffusion ModelShutong Wu, Jiongxiao Wang, Wei Ping, Weili Nie et al.ICLR 2023 · 6 citations
Builds on9
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 1,765 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- DolphinAttack: Inaudible Voice CommandsGuoming Zhang, Chen Yan, Xiaoyu Ji, Tianchen Zhang et al.CCS 2017 · 753 citations
- Hidden Voice CommandsNicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang et al.USENIX Security 2016 · 672 citations
Related papers
- Demystifying Limited Adversarial Transferability in Automatic Speech Recognition SystemsHadi Abdullah, Aditya Karlekar, Vincent Bindschaedler, Patrick TraynorICLR 2022 · 12 citations
- TrojanModel: A Practical Trojan Attack against Automatic Speech Recognition SystemsWei Zong, Yang-Wai Chow, Willy Susilo, Kien Do et al.S&P 2023
- Practical Hidden Voice Attacks against Speech and Speaker Recognition SystemsHadi Abdullah, Washington Garcia, Christian Peeters, Patrick Traynor et al.NDSS 2019 · 180 citations
- Tubes Among Us: Analog Attack on Automatic Speaker IdentificationShimaa Ahmed, Yash Wani, Ali Shahin Shamsabadi, Mohammad Yaghini et al.USENIX Security 2023
- Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation ModelsVyas Raina, Rao Ma, Charles McGhee, Kate M. Knill et al.EMNLP 2024 · 5 citations
