Black-box Adversarial Attacks on Commercial Speech Platforms with Minimal Information
Baolin Zheng, Peipei Jiang, Qian Wang, Qi Li, Chao Shen, Cong Wang, Yunjie Ge, Qingyang Teng, Shenyi Zhang
Abstract
Adversarial attacks against commercial black-box speech platforms, including cloud speech APIs and voice control devices, have received little attention until recent years. Constructing such attacks is difficult mainly due to the unique characteristics of time-domain speech signals and the much more complex architecture of acoustic systems. The current "black-box" attacks all heavily rely on the knowledge of prediction/confidence scores or other probability information to craft effective adversarial examples (AEs), which can be intuitively defended by service providers without returning these messages. In this paper, we take one more step forward and propose two novel adversarial attacks in more practical and rigorous scenarios. For commercial cloud speech APIs, we propose Occam, a decision-only black-box adversarial attack, where only final decisions are available to the adversary. In Occam, we formulate the decision-only AE generation as a discontinuous large-scale global optimization problem, and solve it by adaptively decomposing this complicated problem into a set of sub-problems and cooperatively optimizing each one. Our Occam is a one-size-fits-all approach, which achieves 100% success rates of attacks (SRoA) with an average SNR of 14.23dB, on a wide range of popular speech and speaker recognition APIs, including Google, Alibaba, Microsoft, Tencent, iFlytek, and Jingdong, outperforming the state-of-the-art black-box attacks. For commercial voice control devices, we propose NI-Occam, the first non-interactive physical adversarial attack, where the adversary does not need to query the oracle and has no access to its internal information and training data. We, for the first time, combine adversarial attacks with model inversion attacks, and thus generate the physically-effective audio AEs with high transferability without any interaction with target devices. Our experimental results show that NI-Occam can successfully fool Apple Siri, Microsoft Cortana, Google Assistant, iFlytek and Amazon Echo with an average SRoA of 52% and SNR of 9.65dB, shedding light on non-interactive physical attacks against voice control devices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 703a29b7-c27d-4968-9e85-51e0990aeca6Cited by top-tier papers25
- RULER: discriminative and iterative adversarial training for deep neural network fairnessGuanhong Tao, Weisong Sun, Tingxu Han, Chunrong Fang et al.FSE 2022 · 29 citations
- SPECPATCH: Human-In-The-Loop Adversarial Audio Spectrogram Patch Attack on Speech RecognitionHanqing Guo, Yuanda Wang, Nikolay Ivanov, Li Xiao et al.CCS 2022 · 22 citations
- ALIF: Low-Cost Adversarial Audio Attacks on Black-Box Speech Platforms using Linguistic FeaturesPeng Cheng, Yuwei Wang, Peng Huang, Zhongjie Ba et al.S&P 2024 · 16 citations
- MASTERKEY: Practical Backdoor Attack Against Speaker Verification SystemsHanqing Guo, Xun Chen, Junfeng Guo, Li Xiao et al.MobiCom 2023 · 14 citations
- When Evil Calls: Targeted Adversarial Voice over IP NetworkHan Liu, Zhiyuan Yu, Mingming Zha, XiaoFeng Wang et al.CCS 2022 · 13 citations
Builds on16
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- HopSkipJumpAttack: A Query-Efficient Decision-Based AttackJianbo Chen, Michael I. Jordan, Martin J. WainwrightS&P 2020 · 797 citations
- DolphinAttack: Inaudible Voice CommandsGuoming Zhang, Chen Yan, Xiaoyu Ji, Tianchen Zhang et al.CCS 2017 · 753 citations
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda et al.NeurIPS 2020 · 697 citations
Related papers
- Devil's Whisper: A General Approach for Physical Adversarial Attacks against Commercial Black-box Speech Recognition DevicesYuxuan Chen, Xuejing Yuan, Jiangshan Zhang, Yue Zhao et al.USENIX Security 2020
- QFA2SR: Query-Free Adversarial Transfer Attacks to Speaker Recognition SystemsGuangke Chen, Yedi Zhang, Zhe Zhao, Fu SongUSENIX Security 2023
- KENKU: Towards Efficient and Stealthy Black-box Adversarial Attacks against ASR SystemsXinghui Wu, Shiqing Ma, Chao Shen, Chenhao Lin et al.USENIX Security 2023
- Zero-Query Adversarial Attack on Black-box Automatic Speech Recognition SystemsZheng Fang, Tao Wang, Lingchen Zhao, Shenyi Zhang et al.CCS 2024 · 11 citations
- More Simplicity for Trainers, More Opportunity for Attackers: Black-Box Attacks on Speaker Recognition Systems by Inferring Feature ExtractorYunjie Ge, Pinji Chen, Qian Wang, Lingchen Zhao et al.USENIX Security 2024 · 4 citations
