USENIX Security2020Top-tier venue
Void: A fast and light voice liveness detection system
Muhammad Ejaz Ahmed, Il-Youp Kwak, Jun Ho Huh, Iljoo Kim, Taekkyung Oh, Hyoungshick Kim
Abstract
Due to the open nature of voice assistants' input channels, adversaries could easily record people's use of voice commands, and replay them to spoof voice assistants. To mitigate such spoofing attacks, we present a highly efficient voice liveness detection solution called "Void." Void detects voice spoofing attacks using the differences in spectral power between live-human voices and voices replayed through speakers. In contrast to existing approaches that use multiple deep learning models, and thousands of features, Void uses a single classification model with just 97 features. We used two datasets to evaluate its performance: (1) 255,173 voice samples generated with 120 participants, 15 playback devices and 12 recording devices, and (2) 18,030 publicly available voice samples generated with 42 participants, 26 playback devices and 25 recording devices. Void achieves equal error rate of 0.3% and 11.6% in detecting voice replay attacks for each dataset, respectively. Compared to a state of the art, deep learning-based solution that achieves 7.4% error rate in that public dataset, Void uses 153 times less memory and is about 8 times faster in detection. When combined with a Gaussian Mixture Model that uses Melfrequency cepstral coefficients (MFCC) as classification features -MFCC is already being extracted and used as the main feature in speech recognition services -Void achieves 8.7% error rate on the public dataset. Moreover, Void is resilient against hidden voice command, inaudible voice command, voice synthesis, equalization manipulation attacks, and combining replay attacks with live-human voices achieving about 99.7%, 100%, 90.2%, 86.3%, and 98.2% detection rates for those attacks, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a230373f-6042-40cd-bc41-00e327664fe5Cited by top-tier papers25
- "Hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real WorldEmily Wenger, Max Bronckers, Christian Cianfarani, Jenna Cryan et al.CCS 2021 · 36 citations
- Robust Detection of Machine-induced Audio Attacks in Intelligent Audio Systems with Microphone ArrayZhuohang Li, Cong Shi, Tianfang Zhang, Yi Xie et al.CCS 2021 · 29 citations
- AntiFake: Using Adversarial Audio to Prevent Unauthorized Speech SynthesisZhiyuan Yu, Shixuan Zhai, Ning ZhangCCS 2023 · 29 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- WavoID: Robust and Secure Multi-modal User Identification via mmWave-voice MechanismTiantian Liu, Feng Lin, Chao Wang, Chenhan Xu et al.UIST 2023 · 19 citations
Builds on2
- Hearing Your Voice is Not Enough: An Articulatory Gesture Based Liveness Detection for Voice AuthenticationLinghan Zhang, Sheng Tan, Jie YangCCS 2017 · 212 citations
- VoiceLive: A Phoneme Localization based Liveness Detection for Voice Authentication on SmartphonesLinghan Zhang, Sheng Tan, Jie Yang, Yingying ChenCCS 2016 · 187 citations
Related papers
- VoShield: Voice Liveness Detection with Sound Field DynamicsQiang Yang, Kaiyan Cui, Yuanqing ZhengINFOCOM 2023 · 13 citations
- Learning Normality is Enough: A Software-based Mitigation against Inaudible Voice AttacksXinfeng Li, Xiaoyu Ji, Chen Yan, Chaohao Li et al.USENIX Security 2023
- Unobtrusive Pedestrian Identification by Leveraging Footstep Sounds with Replay ResistanceLong Huang, Chen WangUbiComp 2022 · 8 citations
- When the Differences in Frequency Domain are Compensated: Understanding and Defeating Modulated Replay Attacks on Automatic Speech RecognitionShu Wang, Jiahao Cao, Xu He, Kun Sun et al.CCS 2020 · 34 citations
- SiFDetectCracker: An Adversarial Attack Against Fake Voice Detection Based on Speaker-Irrelative FeaturesXuan Hai, Xin Liu, Yuan Tan, Qingguo ZhouACM MM 2023 · 6 citations
