Robust Detection of Machine-induced Audio Attacks in Intelligent Audio Systems with Microphone Array
Zhuohang Li, Cong Shi, Tianfang Zhang, Yi Xie, Jian Liu, Bo Yuan, Yingying Chen
Abstract
With the popularity of intelligent audio systems in recent years, their vulnerabilities have become an increasing public concern. Existing studies have designed a set of machine-induced audio attacks, such as replay attacks, synthesis attacks, hidden voice commands, inaudible attacks, and audio adversarial examples, which could expose users to serious security and privacy threats. To defend against these attacks, existing efforts have been treating them individually. While they have yielded reasonably good performance in certain cases, they can hardly be combined into an all-in-one solution to be deployed on the audio systems in practice. Additionally, modern intelligent audio devices, such as Amazon Echo and Apple HomePod, usually come equipped with microphone arrays for far-field voice recognition and noise reduction. Existing defense strategies have been focusing on single- and dual-channel audio, while only few studies have explored using multi-channel microphone array for defending specific types of audio attack. Motivated by the lack of systematic research on defending miscellaneous audio attacks and the potential benefits of multi-channel audio, this paper builds a holistic solution for detecting machine-induced audio attacks leveraging multi-channel microphone arrays on modern intelligent audio systems. Specifically, we utilize magnitude and phase spectrograms of multi-channel audio to extract spatial information and leverage a deep learning model to detect the fundamental difference between human speech and adversarial audio generated by the playback machines. Moreover, we adopt an unsupervised domain adaptation training framework to further improve the model's generalizability in new acoustic environments. Evaluation is conducted under various settings on a public multi-channel replay attack dataset and a self-collected multi-channel audio attack dataset involving 5 types of advanced audio attacks. The results show that our method can achieve an equal error rate (EER) as low as 6.6% in detecting a variety of machine-induced attacks. Even in new acoustic environments, our method can still achieve an EER as low as 8.8%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ed4e6e1-319a-4776-9446-4a00e0234f69Cited by top-tier papers12
- mmMIC: Multi-modal Speech Recognition based on mmWave RadarLong Fan, Lei Xie, Xinran Lu, Yi Li et al.INFOCOM 2023 · 41 citations
- SPECPATCH: Human-In-The-Loop Adversarial Audio Spectrogram Patch Attack on Speech RecognitionHanqing Guo, Yuanda Wang, Nikolay Ivanov, Li Xiao et al.CCS 2022 · 22 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- XPorter: A Study of the Multi-Port Charger Security on Privacy Leakage and Voice InjectionTao Ni, Yongliang Chen, Weitao Xu, Lei Xue et al.MobiCom 2023 · 15 citations
- MASTERKEY: Practical Backdoor Attack Against Speaker Verification SystemsHanqing Guo, Xun Chen, Junfeng Guo, Li Xiao et al.MobiCom 2023 · 14 citations
Builds on13
- DolphinAttack: Inaudible Voice CommandsGuoming Zhang, Chen Yan, Xiaoyu Ji, Tianchen Zhang et al.CCS 2017 · 753 citations
- Hidden Voice CommandsNicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang et al.USENIX Security 2016 · 672 citations
- CommanderSong: A Systematic Approach for Practical Adversarial Voice RecognitionXuejing Yuan, Yuxuan Chen, Yue Zhao, Yunhui Long et al.USENIX Security 2018 · 389 citations
- Hearing Your Voice is Not Enough: An Articulatory Gesture Based Liveness Detection for Voice AuthenticationLinghan Zhang, Sheng Tan, Jie YangCCS 2017 · 212 citations
- VoiceLive: A Phoneme Localization based Liveness Detection for Voice Authentication on SmartphonesLinghan Zhang, Sheng Tan, Jie Yang, Yingying ChenCCS 2016 · 187 citations
Related papers
- When the Differences in Frequency Domain are Compensated: Understanding and Defeating Modulated Replay Attacks on Automatic Speech RecognitionShu Wang, Jiahao Cao, Xu He, Kun Sun et al.CCS 2020 · 34 citations
- A Unified Framework for Detecting Audio Adversarial ExamplesXia Du, Chi-Man Pun, Zheng ZhangACM MM 2020 · 18 citations
- Echo: Reverberation-based Fast Black-Box Adversarial Attacks on Intelligent Audio SystemsMeng Xue, Kuang Peng, Xueluan Gong, Qian Zhang et al.UbiComp 2023 · 2 citations
- Voiceprint Mimicry Attack Towards Speaker Verification System in Smart HomeLei Zhang, Yan Meng, Jiahao Yu, Chong Xiang et al.INFOCOM 2020 · 49 citations
- WaveGuard: Understanding and Mitigating Audio Adversarial ExamplesShehzeen Hussain, Paarth Neekhara, Shlomo Dubnov, Julian J. McAuley et al.USENIX Security 2021 · 89 citations
