Hear "No Evil", See "Kenansville"*: Efficient and Transferable Black-Box Attacks on Speech Recognition and Voice Identification Systems
Hadi Abdullah, Muhammad Sajidur Rahman, Washington Garcia, Kevin Warren, Anurag Swarnim Yadav, Tom Shrimpton, Patrick Traynor
摘要
Automatic speech recognition and voice identification systems are being deployed in a wide array of applications, from providing control mechanisms to devices lacking traditional interfaces, to the automatic transcription of conversations and authentication of users. Many of these applications have significant security and privacy considerations. We develop attacks that force mistranscription and misidentification in state of the art systems, with minimal impact on human comprehension. Processing pipelines for modern systems are comprised of signal preprocessing and feature extraction steps, whose output is fed to a machine-learned model. Prior work has focused on the models, using white-box knowledge to tailor model-specific attacks. We focus on the pipeline stages before the models, which (unlike the models) are quite similar across systems. As such, our attacks are black-box and transferable, and demonstrably achieve mistranscription and misidentification rates as high as 100% by modifying only a few frames of audio. We perform a study via Amazon Mechanical Turk demonstrating that there is no statistically significant difference between human perception of regular and perturbed audio. Our findings suggest that models may learn aspects of speech that are generally not perceived by human subjects, but that are crucial for model accuracy. We also find that certain English language phonemes (in particular, vowels) are significantly more susceptible to our attack. We show that the attacks are effective when mounted over cellular networks, where signals are subject to degradation due to transcoding, jitter, and packet loss.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Who is Real Bob? Adversarial Attacks on Speaker Recognition SystemsGuangke Chen, Sen Chen, Lingling Fan, Xiaoning Du 等S&P 2021 · 被引用 239 次
- VoiceBlock: Privacy through Real-Time Adversarial Attacks with Audio-to-Audio ModelsPatrick O'Reilly, Andreas Bugler, Keshav Bhandari, Max Morrison 等NeurIPS 2022 · 被引用 18 次
- When Evil Calls: Targeted Adversarial Voice over IP NetworkHan Liu, Zhiyuan Yu, Mingming Zha, XiaoFeng Wang 等CCS 2022 · 被引用 13 次
- Perception-Aware Attack: Creating Adversarial Music via Reverse-Engineering Human PerceptionRui Duan, Zhe Qu, Shangqing Zhao, Leah Ding 等CCS 2022 · 被引用 8 次
- Defending against Adversarial Audio via Diffusion ModelShutong Wu, Jiongxiao Wang, Wei Ping, Weili Nie 等ICLR 2023 · 被引用 6 次
它引用的顶会 Paper9
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 被引用 1,765 次
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- DolphinAttack: Inaudible Voice CommandsGuoming Zhang, Chen Yan, Xiaoyu Ji, Tianchen Zhang 等CCS 2017 · 被引用 753 次
- Hidden Voice CommandsNicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang 等USENIX Security 2016 · 被引用 672 次
相关 Paper
- Demystifying Limited Adversarial Transferability in Automatic Speech Recognition SystemsHadi Abdullah, Aditya Karlekar, Vincent Bindschaedler, Patrick TraynorICLR 2022 · 被引用 12 次
- TrojanModel: A Practical Trojan Attack against Automatic Speech Recognition SystemsWei Zong, Yang-Wai Chow, Willy Susilo, Kien Do 等S&P 2023
- Practical Hidden Voice Attacks against Speech and Speaker Recognition SystemsHadi Abdullah, Washington Garcia, Christian Peeters, Patrick Traynor 等NDSS 2019 · 被引用 180 次
- Tubes Among Us: Analog Attack on Automatic Speaker IdentificationShimaa Ahmed, Yash Wani, Ali Shahin Shamsabadi, Mohammad Yaghini 等USENIX Security 2023
- Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation ModelsVyas Raina, Rao Ma, Charles McGhee, Kate M. Knill 等EMNLP 2024 · 被引用 5 次
