Breaking Security-Critical Voice Authentication
Andre Kassis, Urs Hengartner
Abstract
Voice authentication (VA) has recently become an integral part in numerous security-critical operations, such as bank transactions and call center conversations. The vulnerability of automatic speaker verification systems (ASVs) to spoofing attacks instigated the development of countermeasures (CMs), whose task is to differentiate between bonafide and spoofed speech. Together, ASVs and CMs form today’s VA systems and are being advertised as an impregnable access control mechanism. We develop the first practical attack on spoofing countermeasures, and demonstrate how a malicious actor may efficiently craft audio samples against these defenses. Previous adversarial attacks against VA have been mainly designed for the whitebox scenario, which assumes knowledge of the system’s internals, or requires large query and time budgets to launch target-specific attacks. When attacking a security-critical system, these assumptions do not hold. Our attack, on the other hand, targets common points of failure that all spoofing countermeasures share, making it real-time, model-agnostic, and completely blackbox without the need to interact with the target to craft the attack samples. The key message from our work is that CMs mistakenly learn to distinguish between spoofed and bonafide audio based on cues that are easily identifiable and forgeable. The effects of our attack are subtle enough to guarantee that these adversarial samples can still bypass the ASV as well and preserve their original textual contents. These properties combined make for a powerful attack that can bypass security-critical VA in its strictest form, yielding success rates of up to 99% with only 6 attempts. Finally, we perform the first targeted, over-telephony-network attack on CMs, bypassing several known challenges and enabling a variety of potential threats, given the increased use of voice biometrics in call centers. Our results call into question the security of modern VA systems and urge users to rethink their trust in them, in light of the real threat of attackers bypassing these measures to gain access to their most valuable resources.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18d56cbc-900e-4174-8868-6dcdf49df146Cited by top-tier papers4
- SiFMimicEvader: Evading Fake Voice Detection with Adversarial Neural Mimicry AttacksXuan Hai, Xin Liu, Zihao Zhang, Ziyao Yu et al.ACM MM 2025
- Open Sesame! On the Security and Memorability of Verbal PasswordsEunsoo Kim, Kiho Lee, Doowon Kim, Hyoungshick KimS&P 2025
- UnMarker: A Universal Attack on Defensive Image WatermarkingAndre Kassis, Urs HengartnerS&P 2025
- What's the Real: A Novel Design Philosophy for Robust AI-Synthesized Voice DetectionXuan Hai, Xin Liu, Yuan Tan, Gang Liu et al.ACM MM 2024
Builds on14
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- CommanderSong: A Systematic Approach for Practical Adversarial Voice RecognitionXuejing Yuan, Yuxuan Chen, Yue Zhao, Yunhui Long et al.USENIX Security 2018 · 389 citations
- Who is Real Bob? Adversarial Attacks on Speaker Recognition SystemsGuangke Chen, Sen Chen, Lingling Fan, Xiaoning Du et al.S&P 2021 · 239 citations
- Hearing Your Voice is Not Enough: An Articulatory Gesture Based Liveness Detection for Voice AuthenticationLinghan Zhang, Sheng Tan, Jie YangCCS 2017 · 212 citations
- VoiceLive: A Phoneme Localization based Liveness Detection for Voice Authentication on SmartphonesLinghan Zhang, Sheng Tan, Jie Yang, Yingying ChenCCS 2016 · 187 citations
Related papers
- Voiceprint Mimicry Attack Towards Speaker Verification System in Smart HomeLei Zhang, Yan Meng, Jiahao Yu, Chong Xiang et al.INFOCOM 2020 · 49 citations
- SiFDetectCracker: An Adversarial Attack Against Fake Voice Detection Based on Speaker-Irrelative FeaturesXuan Hai, Xin Liu, Yuan Tan, Qingguo ZhouACM MM 2023 · 6 citations
- When Evil Calls: Targeted Adversarial Voice over IP NetworkHan Liu, Zhiyuan Yu, Mingming Zha, XiaoFeng Wang et al.CCS 2022 · 13 citations
- Tubes Among Us: Analog Attack on Automatic Speaker IdentificationShimaa Ahmed, Yash Wani, Ali Shahin Shamsabadi, Mohammad Yaghini et al.USENIX Security 2023
- SMACK: Semantically Meaningful Adversarial Audio AttackZhiyuan Yu, Yuanhaur Chang, Ning Zhang, Chaowei XiaoUSENIX Security 2023
