USENIX Security2023Top-tier venue
Learning Normality is Enough: A Software-based Mitigation against Inaudible Voice Attacks
Xinfeng Li, Xiaoyu Ji, Chen Yan, Chaohao Li, Yichen Li, Zhenning Zhang, Wenyuan Xu
Abstract
Inaudible voice attacks silently inject malicious voice commands into voice assistants to manipulate voice-controlled devices such as smart speakers. To alleviate such threats for both existing and future devices, this paper proposes NormDetect, a software-based mitigation that can be instantly applied to a wide range of devices without requiring any hardware modification. To overcome the challenge that the attack patterns vary between devices, we design a universal detection model that does not rely on audio features or samples derived from specific devices. Unlike existing studies' supervised learning approach, we adopt unsupervised learning inspired by anomaly detection. Though the patterns of inaudible voice attacks are diverse, we find that benign audios share similar patterns in the time-frequency domain. Therefore, we can detect the attacks (the anomaly) by learning the patterns of benign audios (the normality). NormDetect maps spectrum features to a low-dimensional space, performs similarity queries, and replaces them with the standard feature embeddings for spectrum reconstruction. This results in a more significant reconstruction error for attacks than normality. Evaluation based on the 383,320 test samples we collected from 24 smart devices shows an average AUC of 99.48% and EER of 2.23%, suggesting the effectiveness of NormDetect in detecting inaudible voice attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2967b40-86e7-4c62-a8fe-b3b706f49e58Cited by top-tier papers7
- Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMsZijian Ling, Pingyi Hu, Xiuyong Gao, Xiaojing Ma et al.USENIX Security 2026 · 185 citations
- SafeEar: Content Privacy-Preserving Audio Deepfake DetectionXinfeng Li, Kai Li, Yifan Zheng, Chen Yan et al.CCS 2024 · 26 citations
- AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language ModelsKai Li, Can Shen, Yile Liu, Jirui Han et al.ICLR 2026 · 17 citations
- Unveiling the Hunter-Gatherers: Exploring Threat Hunting Practices and Challenges in Cyber DefensePriyanka Badva, Kopo M. Ramokapane, Eleonora Pantano, Awais RashidUSENIX Security 2024 · 13 citations
- MicGuard: A Comprehensive Detection System against Out-of-band Injection Attacks for Different Level Microphone-based DevicesTiantian Liu, Feng Lin, Zhongjie Ba, Li Lu et al.USENIX Security 2024 · 4 citations
Builds on8
- DolphinAttack: Inaudible Voice CommandsGuoming Zhang, Chen Yan, Xiaoyu Ji, Tianchen Zhang et al.CCS 2017 · 753 citations
- SoK: The Faults in our ASRs: An Overview of Attacks against Automatic Speech Recognition and Speaker Identification SystemsHadi Abdullah, Kevin Warren, Vincent Bindschaedler, Nicolas Papernot et al.S&P 2021 · 145 citations
- Robust Detection of Machine-induced Audio Attacks in Intelligent Audio Systems with Microphone ArrayZhuohang Li, Cong Shi, Tianfang Zhang, Yi Xie et al.CCS 2021 · 29 citations
- CapSpeaker: Injecting Voices to Microphones via CapacitorsXiaoyu Ji, Juchuan Zhang, Shui Jiang, Jishen Li et al.CCS 2021 · 21 citations
- SurfingAttack: Interactive Hidden Attack on Voice Assistants Using Ultrasonic Guided WavesQiben Yan, Kehai Liu, Qin Zhou, Hanqing Guo et al.NDSS 2020
Related papers
- EarArray: Defending against DolphinAttack via Acoustic AttenuationGuoming Zhang, Xiaoyu Ji, Xinfeng Li, Gang Qu et al.NDSS 2021
- Inducing Wireless Chargers to Voice Out for Inaudible Command AttacksDonghui Dai, Zhenlin An, Lei YangS&P 2023
- CommanderSong: A Systematic Approach for Practical Adversarial Voice RecognitionXuejing Yuan, Yuxuan Chen, Yue Zhao, Yunhui Long et al.USENIX Security 2018 · 389 citations
- GhostTalk: Interactive Attack on Smartphone Voice System Through Power LineYuanda Wang, Hanqing Guo, Qiben YanNDSS 2022
- When the Differences in Frequency Domain are Compensated: Understanding and Defeating Modulated Replay Attacks on Automatic Speech RecognitionShu Wang, Jiahao Cao, Xu He, Kun Sun et al.CCS 2020 · 34 citations
